AI and Data Scientist
From data to models you can defend.
The complete data-science path: math that matters, the Python stack, data wrangling and EDA, classical machine learning done honestly, deep learning, transformers and LLMs, and the portfolio that proves it.
Saved on this device - no account needed.
- 1
Math That Matters
Just enough math, deeply understood.
3 weeks0/3Vectors, matrices and dot products - the notation every model description is written in.
Free resources
Derivatives and gradients as 'which way is downhill' - the idea behind all training.
Free resources
Random variables, distributions, conditional probability and Bayes - the language of uncertainty.
Free resources
- 2
Python & The PyData Stack
Make Python your native tool for data.
3–4 weeks0/4Functions, comprehensions, error handling and environments - clean code before clever models.
Free resources
Arrays, broadcasting and vectorised thinking - the substrate every ML library builds on.
Free resources
Loading, cleaning, joining and reshaping real datasets - where data scientists spend most hours.
Free resources
Notebooks for exploration, scripts for pipelines, seeds and environments for results others can rerun.
Free resources
- 3
Data Wrangling & EDA
Interrogate a dataset before modelling it.
2–3 weeks0/4Missing values, duplicates, outliers and type fixes - with the judgement calls documented.
Free resources
Distributions, relationships and anomalies via plots and groupbys - the stage that decides whether modelling is even warranted.
Free resources
matplotlib and seaborn fluency - plots that expose bad data before it fools your model.
Free resources
Joins, aggregation and window functions - most industry data still lives behind a SQL interface.
Free resources
Build: EDA case study
Pick a public dataset; deliver a notebook with cleaning, five insights and honest caveats.
- 4
Statistics for Inference
Reason under uncertainty like a scientist.
2–3 weeks0/3Sampling distributions, standard errors and confidence intervals - what your point estimate is hiding.
Nulls, p-values, power and multiple-comparison traps; why p-hacking ruins careers.
Free resources
Randomisation, metrics and guardrails - designing A/B tests whose answers you can trust.
- 5
Classical Machine Learning
Model tabular data - still most industry ML.
4–5 weeks0/7Linear and regularised regression: coefficients, residuals and interpretation you can defend.
Free resources
Logistic regression, decision trees, random forests, KNN and Naive Bayes - the standard cast and when each shines.
Free resources
XGBoost/LightGBM - the tabular-data workhorses that win competitions and quietly run industry.
Free resources
k-means, hierarchical clustering and PCA - structure-finding when there are no labels.
Train/test discipline, cross-validation and the leakage patterns behind most 'amazing' models.
Precision/recall, ROC-AUC, calibration and cost-sensitive choices - the metric the business cares about, not the default.
Encoding, scaling, interactions and domain features - better features beat fancier models.
Free resources
Build: End-to-end prediction model
A churn or price model: EDA, features, model comparison, error analysis and a business-facing summary.
- 6
Deep Learning
Neural networks from intuition to practice.
4–5 weeks0/4Layers, activations, loss and backprop intuition - training as gradient descent on purpose.
Free resources
Tensors, autograd, DataLoaders and the training loop - the research-to-production standard.
Free resources
Convolutions for images, recurrence for sequences - the classic architectures and the intuitions that survive them.
Overfitting, regularisation, learning-rate schedules and transfer learning - the knobs that make models actually converge.
- 7
Transformers & LLMs
The modern layer on top of the fundamentals.
3–4 weeks0/4Self-attention, positional encoding and pretraining - what changed in 2017 and why it took over every modality.
Free resources
Loading pretrained models, pipelines and fine-tuning on your own data.
Free resources
Embeddings, RAG and prompt-based classification - combining foundation models with your own data products.
Experiment tracking, model registries and monitoring drift - models as living software, not one-off notebooks.
Free resources
- 8
Portfolio & Proof
Evidence that survives recruiter scrutiny.
2 weeks0/3Two deep projects with written conclusions beat ten Kaggle notebooks; publish code and findings.
Free resources
Kaggle competitions and study groups - feedback loops that compound faster than solo study.
Free resources
Take the OneRoadmap AI & Data Scientist certification for verified, shareable proof.
Free resources