Structured data
Regression
price, lifetime value, delivery time, energy load
A continuous target per row.
Maximum likelihood
Linear algebra
Bias & variance
Hypothesis testing
Foundations
Least squares is maximum likelihood under Gaussian noise. Regularisation is the bias-variance tradeoff made explicit.
Frame the problem
Frame
Decide whether you need the mean, a quantile, or an interval. Errors are rarely symmetric; say which direction hurts.
Sourcing & signal
Exploring the data
Hunting for leakage
Cleaning the data
Data & labels
Look at the target distribution; long tails usually want a log transform. Leakage hides in fields filled in after the outcome: final invoice, actual delivery time.
Encode & scale
Create features
Reduce & select
Represent
Split the data
Cross-validation
Split
Group by the entity if it repeats. Temporal if prices drift.
Dumb baseline
Linear models & GLMs
Regularised linear
K-nearest neighbors
Gradient boosting
Kernel methods
Model
Predict the mean, then a linear model, and read its coefficients. Ridge when features are many and correlated, lasso when you want fewer. Then gradient boosting. Kernel methods or GPs when data is small and you need uncertainty.
Loss functions
Hyperparameter tuning
Over- & underfitting
Ensembling
Train
Pick the loss to match the framing: MAE for medians, quantile loss for intervals, Huber for outliers.
Regression metrics
Interpretability
Error analysis & significance
Evaluate offline
RMSE and MAE together. Residuals against every feature. Slice by price band.
Evaluate online
Nothing unusual here.
Features at serving time
Serving & release
Monitor & retrain
Ship & monitor
Press enter or space to select a node. You can then use the arrow keys to move the node around. Press delete to remove it and escape to cancel.
Press enter or space to select an edge. You can then press delete to remove it or escape to cancel.