[son of anton]

Lectures

Every topic in the map, grouped by the step it belongs to. Lectures open in a new tab.

9 lectures over 119 topics

1.Foundations

The theory the other steps lean on. Skim it first; come back when a step needs it.

  • Probability & Bayes
    Distributions, conditional probability, Bayes' theorem.no lecture yet
  • Maximum likelihood
    Fitting parameters by maximising the likelihood of the data.no lecture yet
  • Odds & log-odds
    Odds, log-odds, odds ratios.no lecture yet
  • Information theory
    Entropy, cross-entropy, KL divergence, mutual information.no lecture yet
  • Bias & variance
    Why models under- and over-fit.no lecture yet
  • Hypothesis testing
    p-values, confidence intervals, power.no lecture yet
  • Linear algebra
    Vectors, matrices, projections, eigendecomposition, SVD.no lecture yet
  • Calculus & gradients
    Derivatives, chain rule, gradients.no lecture yet
  • Optimisation
    Gradient descent, SGD, convexity, learning rates.no lecture yet
  • Potential outcomes
    Counterfactuals and confounding.no lecture yet
  • MDPs & Bellman
    States, actions, rewards, value functions.no lecture yet

2.Frame

Decide what is being predicted or decided, for whom, and how success is measured.

  • Frame the problem
    Do we need ML? Target, unit of prediction, horizon, metric, baseline, task type.no lecture yet

3.Data & labels

Get the rows, understand them, and make sure the labels mean what you think.

  • Sourcing & signal
    Where the data comes from; enough rows, enough positives, label quality.no lecture yet
  • Exploring the data
    Distributions, ranges, units, target balance, correlations and redundancies, statistical tests of feature → target relationships, missingness patterns.no lecture yet
  • Hunting for leakage
    Features that encode the answer; anything computed after prediction time.no lecture yet
  • Labeling & data collection
    Annotation and agreement, weak supervision, active learning, semi-supervised.no lecture yet
  • Cleaning the data
    Missing values: drop if rare, impute, model-based imputation, sentinel + missing flag, don't over-impute. Outliers: genuine extremes vs data-entry errors; cap, winsorise, log-transform. Consistency: types, units, timezones, encodings, deduplication.no lecture yet

4.Represent

Turn whatever you have into vectors a model can use.

  • Encode & scale
    One-hot, ordinal, target encoding, hashing; when to scale.no lecture yet
  • Create features
    Ratios, interactions, date parts, rolling aggregates, lags.no lecture yet
  • Modalities → vectors
    Text, images, audio, graphs into vectors via TF-IDF or pretrained encoders.no lecture yet
  • Reduce & select
    Drop redundant or leaky features; importance, mutual information, PCA.no lecture yet
  • Embedding layers
    Learned dense vectors for high-cardinality categoricals.no lecture yet
  • word2vec & GloVe
    Embeddings as learned similarity.no lecture yet
  • Tokenization
    BPE, vocabularies, what a token is.no lecture yet

5.Split

Hold out data the way production will hold out the future.

  • Split the data
    Train/val/test; random, stratified, grouped.no lecture yet
  • Temporal splits & backtesting
    Rolling and expanding windows; never leak the future.no lecture yet
  • Cross-validation
    K-fold and variants; nested CV.no lecture yet

6.Model

Start with the dumbest thing that could work, then climb only as far as the evaluation demands.

Baselines & linear models

  • Dumb baseline
    Majority class, mean, last value, popularity, rules.no lecture yet
  • Linear models & GLMs
    Linear regression; GLMs with link functions, deviance and saturated models; Poisson/gamma; quantile and Huber; GAMs.no lecture yet
  • Regularised linear
    Ridge, lasso, elastic net.no lecture yet
  • Binary and softmax; one-vs-rest.

Neighbours, kernels & probabilistic

  • Instance-based prediction.
  • Plus LDA and QDA.
  • Kernel methods
    SVM, SVR, kernel ridge, Gaussian processes.no lecture yet
  • Graphical models
    HMM, CRF, Bayesian networks, EM.no lecture yet
  • Survival analysis
    Kaplan-Meier, Cox, random survival forests, concordance.no lecture yet

Trees & ensembles

  • Greedy splitting, impurity, pruning.
  • Bagging & random forests
    Bootstrap aggregation, random forest, ExtraTrees.no lecture yet
  • GBM, XGBoost/LightGBM/CatBoost.
  • Networks on tabular data
    MLP with embeddings, TabNet, FT-Transformer; when they beat boosting.no lecture yet

Unsupervised

  • Clustering
    K-means, DBSCAN/HDBSCAN, hierarchical, spectral.no lecture yet
  • Gaussian mixtures & EM
    Soft clustering via EM.no lecture yet
  • Dimensionality reduction
    PCA/SVD, NMF, t-SNE/UMAP, topic models.no lecture yet
  • Anomaly detection
    Isolation forest, LOF, one-class SVM, reconstruction error.no lecture yet
  • Association rules
    Apriori, FP-Growth; market basket analysis.no lecture yet

Neural networks

  • MLP & backprop
    The multilayer perceptron and how gradients flow.no lecture yet
  • Autoencoders
    Compression by reconstruction.no lecture yet
  • Convolutional networks
    Convolution, pooling, ResNets, augmentation, transfer.no lecture yet
  • Detection & segmentation
    YOLO, U-Net, Mask R-CNN.no lecture yet
  • Vision transformers & CLIP
    Patches as tokens; vision-language embeddings.no lecture yet
  • RNN / LSTM / GRU
    Recurrent sequence models; TCNs.no lecture yet
  • Attention & transformers
    Attention as soft lookup; the transformer.no lecture yet
  • Transformer internals
    Attention variants, positional encodings, KV cache.no lecture yet
  • Seq2seq
    Encoder-decoder.no lecture yet
  • Sequence labeling & NER
    BiLSTM-CRF, token classification with a fine-tuned encoder.no lecture yet
  • Graph neural networks
    GCN, GraphSAGE; node classification and link prediction.no lecture yet

Sequences, time series & audio

  • Classical time series
    ARIMA/SARIMA, ETS, VAR, state-space, Prophet.no lecture yet
  • Boosting on lag features
    One global gradient-boosted model across many series.no lecture yet
  • Deep forecasting
    DeepAR, TFT, N-BEATS, foundation models.no lecture yet
  • Audio networks
    Spectrograms, MFCCs, 1D CNNs and RNNs.no lecture yet
  • CTC & speech recognition
    Alignment-free sequence loss.no lecture yet
  • Pretrained audio encoders
    wav2vec, Whisper.no lecture yet

Recommenders, search & matching

  • Collaborative filtering
    Item-item and user-user neighbourhoods.no lecture yet
  • Matrix factorization
    SVD, ALS, BPR.no lecture yet
  • Factorization machines
    Pairwise interactions via factorized weights.no lecture yet
  • TF-IDF + linear
    Bag-of-words baseline.no lecture yet
  • BM25 & lexical search
    Term weighting for retrieval.no lecture yet
  • Learning to rank
    Pointwise, pairwise, listwise; LambdaMART.no lecture yet
  • Metric learning
    Siamese, triplet, contrastive losses; negative sampling.no lecture yet
  • Two-tower retrieval & ANN
    Dual encoders; HNSW, IVF-PQ, FAISS.no lecture yet
  • Deep rankers
    Wide & Deep, DeepFM, DLRM, DCN, MMoE.no lecture yet
  • Bi- vs cross-encoders
    Dense retrieval plus re-ranking.no lecture yet
  • Sequential recommenders
    GRU4Rec, SASRec, BERT4Rec.no lecture yet
  • Graph recommenders
    LightGCN, PinSage.no lecture yet
  • The funnel & cold start
    Candidate generation → ranking → re-ranking; fallbacks.no lecture yet
  • Entity resolution
    Blocking, pairwise matching, transitive closure over the match graph.no lecture yet

Using pretrained language models

  • The LLM ladder
    Prompting → RAG → fine-tuning → pretraining.no lecture yet
  • Structured extraction with LLMs
    Schemas, function calling, validation, routing low confidence to humans.no lecture yet

Generative

  • Variational autoencoders
    Latent-variable generation.no lecture yet
  • GANs
    Adversarial training.no lecture yet
  • Normalizing flows
    Invertible maps.no lecture yet
  • Diffusion models
    Iterative denoising.no lecture yet

Decisions: bandits, RL, causal

  • Bandits
    Epsilon-greedy, UCB, Thompson sampling, LinUCB.no lecture yet
  • Online learning
    Streaming SGD, FTRL.no lecture yet
  • Causal inference
    Propensity, matching, IPW, difference-in-differences.no lecture yet
  • Uplift modelling
    Uplift trees, meta-learners, Qini curves.no lecture yet
  • Attribution & marketing mix models
    Multi-touch attribution; aggregate regression with adstock and saturation.no lecture yet
  • Value-based RL
    SARSA, Q-learning, DQN.no lecture yet
  • Policy-based RL
    Policy gradient, actor-critic, PPO.no lecture yet
  • Offline & model-based RL
    Learning from logs; world models.no lecture yet

7.Train

Fit it, tune it, and stop it from memorising.

Fitting classical models

  • Loss functions
    Cross-entropy, hinge; MSE, MAE, Huber, quantile; pairwise and listwise.no lecture yet
  • Hyperparameter tuning
    Grid, random, Bayesian.no lecture yet
  • Handle imbalance
    Class weights, resampling, thresholds.no lecture yet
  • Over- & underfitting
    Regularisation, early stopping, learning curves.no lecture yet
  • Ensembling
    Voting, bagging vs boosting, stacking.no lecture yet
  • Calibration
    Platt, isotonic, reliability curves; recalibrating after downsampling.no lecture yet

Training neural networks

  • Training craft
    Init, normalisation, dropout, schedules, clipping, debugging loss curves.no lecture yet
  • Fine-tuning a pretrained model
    Freezing, unfreezing, discriminative learning rates.no lecture yet

Pretraining & adaptation

  • Next-token pretraining
    Train a nano-scale GPT.no lecture yet
  • Masked & contrastive pretraining
    BERT-style objectives.no lecture yet
  • Fine-tuning & LoRA
    Full vs parameter-efficient.no lecture yet
  • Distillation & domain adaptation
    Small models from big ones.no lecture yet
  • RLHF & DPO
    Alignment from preferences.no lecture yet

8.Evaluate offline

Measure it on held-out data, then find out where it is wrong.

Metrics by task

  • Accuracy, precision/recall/F1, sensitivity and specificity, ROC-AUC, PR-AUC, log loss, confusion matrix.
  • RMSE, MAE, MAPE, R².
  • Forecasting metrics
    MAPE, sMAPE, MASE vs naive.no lecture yet
  • NDCG, MAP, MRR, recall@k, coverage.
  • Cosine vs dot, anisotropy, recall@k.
  • Evaluating without labels
    Silhouette, ARI, precision at alert budget.no lecture yet
  • Evaluating generative output
    Perplexity, benchmarks, LLM-as-judge, preference.no lecture yet

Understanding errors

  • Interpretability
    Global: permutation importance, partial dependence. Local: SHAP, LIME, counterfactuals, reason codes.no lecture yet
  • Error analysis & significance
    Slices, fairness, robustness, significance of comparisons.no lecture yet

9.Evaluate online

Find out whether the offline number meant anything.

  • Offline / online gap
    Why offline metrics disagree with production.no lecture yet
  • A/B testing & interleaving
    Randomisation, power, guardrails, novelty; interleaving for ranking.no lecture yet
  • Off-policy evaluation
    IPS, doubly robust.no lecture yet

10.Ship & monitor

Serve it, watch it, and know when to retrain.

  • Features at serving time
    Same transform online; pipelines and feature stores.no lecture yet
  • Serving & release
    Batch/realtime/edge; latency, throughput and cost budgets; experiment tracking (seeds, data versioning, MLflow/W&B); model registry, versioning, rollback; shadow mode, canary.no lecture yet
  • Inference optimisation
    Quantisation, speculative decoding, batching, caching.no lecture yet
  • Monitor & retrain
    Log predictions; data drift, concept drift, performance decay; retraining triggers; feedback-loop awareness.no lecture yet