Decide & act
Budget-constrained optimisation
bid shading, campaign pacing, inventory allocation
Spend a budget across auctions, channels, or inventory to maximise an outcome.
MDPs & Bellman
Probability & Bayes
Optimisation
Foundations
Bellman equations frame the sequential problem; constrained optimisation frames the budget.
Frame the problem
Frame
A sequential decision under a budget constraint. Define the reward, the horizon, and the constraint.
Sourcing & signal
Hunting for leakage
Data & labels
Auction logs with what you bid, whether you won, and what it cost. Censored: you do not see the clearing price when you lose.
Create features
Represent
State features: budget remaining, time remaining, pacing so far.
Temporal splits & backtesting
Split
Temporal, by campaign.
Dumb baseline
Logistic regression
Gradient boosting
Calibration
Bandits
Contextual bandits for allocation across a small action set with fast feedback.
Bandits
Online learning
Reinforcement learning
Only when actions have long consequences. Train offline from logs first; online RL against a live auction is expensive.
Value-based RL
Policy-based RL
Offline & model-based RL
Model
Fixed bids or a proportional pacing rule; simple controllers are hard to beat. Bidding needs calibrated value and win-rate predictions, which are their own playbooks.
Train
Nothing unusual here.
Off-policy evaluation
Evaluate offline
Off-policy evaluation from logs before any live test.
A/B testing & interleaving
Evaluate online
A budget-split A/B.
Serving & release
Monitor & retrain
Ship & monitor
Guardrails: spend caps and fallbacks to the baseline controller. Monitor constraint violations, not just reward.
Press enter or space to select a node. You can then use the arrow keys to move the node around. Press delete to remove it and escape to cancel.
Press enter or space to select an edge. You can then press delete to remove it or escape to cancel.