If you like building the pipes that move and clean data so everyone else can trust it, this role is for you.
Builds the pipelines that move and clean data so analysts and models can use it. Pipelines, SQL tooling and orchestration are mostly written in Python.
Take the lessons top to bottom. Each one opens with code you run, and ends with graded exercises that check your answer.
16 lessons every path shares. Skip the ones you already know.
Write your first Python program and print to the screen.
Numbers, strings, booleans — the building blocks.
Arithmetic, powers, the math module — Python as a calculator.
Slicing, methods, and f-strings — text manipulation done right.
Make decisions and repeat yourself — but only on purpose.
Loop until a condition fails. Bail out or skip with break/continue.
Collections of things, and how to walk through them.
Carve up lists and strings with [start:stop:step].
Two more collections: tuples are fixed, sets are unique.
Key → value pairs. The most important data structure in Python.
Empty things are False. `is` vs `==`. The infamous None.
The Pythonic ways to iterate — count, pair, sequence.
Bundle up logic so you can reuse it.
Lists of dicts, dicts of lists — modelling real-world data.
Filter and transform in one line — the Pythonic way.
Catch what could go wrong, recover gracefully.
15 lessons specific to this role.
Lazy iteration with `yield` — handle huge sequences in O(1) memory.
Parse and produce JSON — the universal data format.
Today, deltas, parsing, formatting — without losing your mind.
A real database in a single file. Perfect for learning SQL end-to-end.
INNER, LEFT, RIGHT, FULL — with a concrete users-and-orders example.
SQLAlchemy Core tutorial: write database SQL with type-safe Python expressions, the foundation beneath the ORM.
Declare Python classes, get automatic CRUD, relationships, and lazy loading.
Python database migrations with Alembic and SQLAlchemy: evolve schemas safely with revisions, autogenerate, and an audit trail.
DataFrames in your browser. Auto-installs pandas on first run.
Real datasets arrive dirty: missing cells, wrong dtypes, duplicate rows. Fix them with isna/fillna/dropna, to_numeric, rename and drop_duplicates.
Combine dataframes like SQL joins — with inner, left, outer, and validation.
Pandas time series tutorial: use DatetimeIndex, resampling, rolling windows, and grouped time-based analysis.
Combine CSV parsing, cleaning, dataclasses and aggregation into one small end-to-end pipeline.
Beyond print() — structured logs, levels, handlers, and what real backends actually emit.
Measure before you optimize — cProfile, timeit, and the shapes real bottlenecks take.
Projects and workspace templates that fit this role. Publish the result to your portfolio page.
Finishing this path gives you the Python part of the job, with graded exercises and projects you can show. It is not, by itself, a qualification for the role.