Data Analysis · pandas Core · lesson 8 of 12
Missing data
about 15 minutes · free · runs in your browser
The gaps are the job
Real data has holes. pandas marks them NaN, and the first thing to understand is
that NaN is not equal to anything, including itself:
float("nan") == float("nan") # False
So you never test for it with ==. Use the methods:
df.isna() # a boolean table
df["temp"].isna().sum() # how many gaps in that column
df.dropna() # drop any row with a gap
df.dropna(subset=["temp"]) # only where temp is missing
df["temp"].fillna(0) # replace them
df["temp"].fillna(df["temp"].mean()) # a common choice
Aggregations skip NaN by default, which is usually what you want but occasionally hides a problem: a mean over mostly-missing data is still a number.
Your turn: count the gaps, then produce a version of the table where missing temperatures are replaced by the column's mean, rounded to one decimal.
You start from this, and edit it in the browser:
import pandas as pd
import numpy as np
df = pd.DataFrame({
"city": ["London", "Lisbon", "Oslo", "Tokyo"],
"temp": [14, np.nan, 3, np.nan],
})
# Set missing_count and filled (a DataFrame with the gaps replaced).