You’ll be able to
- Load CSV and Excel files into pandas DataFrames
- Filter rows and select columns with boolean indexing
- Group data with
.groupby()and aggregate with.sum(),.mean() - Handle missing values with
.dropna()and.fillna()
Loading lesson…
Why this matters
Pandas is the lingua franca of data engineering, ML feature pipelines, and analyst-facing ETL — every Airflow DAG and Jupyter notebook eventually touches a DataFrame. Staff engineers avoid .apply() and .iterrows() in favor of vectorized column ops, groupby().agg(), and merge with explicit how= and validate= to catch join-cardinality bugs early.
Common pitfalls
- Chained assignment like
df[df.x > 0]['y'] = 1— triggersSettingWithCopyWarning; use.loc[mask, 'y'] = 1. - Merging without
validate='one_to_one'— silent row explosions from duplicate keys ruin downstream counts. - Storing strings as
objectdtype — 10× more memory thancategoryfor low-cardinality columns like brand or region.