If you like asking ‘what will happen next?’, testing ideas against real data, and explaining a model to people who never saw the code, this role is for you.
Uses data and models to predict outcomes and explain what drives them. Python is the default language of data science: pandas for data, scikit-learn for models.
Take the lessons top to bottom. Each one opens with code you run, and ends with graded exercises that check your answer.
16 lessons every path shares. Skip the ones you already know.
Write your first Python program and print to the screen.
Numbers, strings, booleans — the building blocks.
Arithmetic, powers, the math module — Python as a calculator.
Slicing, methods, and f-strings — text manipulation done right.
Make decisions and repeat yourself — but only on purpose.
Loop until a condition fails. Bail out or skip with break/continue.
Collections of things, and how to walk through them.
Carve up lists and strings with [start:stop:step].
Two more collections: tuples are fixed, sets are unique.
Key → value pairs. The most important data structure in Python.
Empty things are False. `is` vs `==`. The infamous None.
The Pythonic ways to iterate — count, pair, sequence.
Bundle up logic so you can reuse it.
Lists of dicts, dicts of lists — modelling real-world data.
Filter and transform in one line — the Pythonic way.
Catch what could go wrong, recover gracefully.
16 lessons specific to this role.
Bundle state and behavior — the basics of OOP in Python.
Parse and produce JSON — the universal data format.
Vectorized math in Python. Auto-installs NumPy in your browser on first run.
Replace Python loops with array math — 10–100× speedups for numerical work.
DataFrames in your browser. Auto-installs pandas on first run.
Real datasets arrive dirty: missing cells, wrong dtypes, duplicate rows. Fix them with isna/fillna/dropna, to_numeric, rename and drop_duplicates.
Split-apply-combine — the pattern behind every real data analysis.
Combine dataframes like SQL joins — with inner, left, outer, and validation.
Pandas time series tutorial: use DatetimeIndex, resampling, rolling windows, and grouped time-based analysis.
Real charts that ship — line, bar, distribution, heatmap.
Train a real scikit-learn model in your browser and make a prediction.
Why you must hold data back, and how to measure honest accuracy.
Fit, predict, and score a regression model with R² and MAE.
Predict yes/no outcomes with logistic regression and read accuracy.
Train a tree you can read, and see which features actually matter.
See overfitting happen, then read a confusion matrix like a pro.
Projects and workspace templates that fit this role. Publish the result to your portfolio page.
Finishing this path gives you the Python part of the job, with graded exercises and projects you can show. It is not, by itself, a qualification for the role.