[son of anton]
Decide & act

Budget-constrained optimisation

bid shading, campaign pacing, inventory allocation

Spend a budget across auctions, channels, or inventory to maximise an outcome.

MDPs & Bellman
Probability & Bayes
Optimisation
Foundations
Bellman equations frame the sequential problem; constrained optimisation frames the budget.
Frame the problem
Frame
A sequential decision under a budget constraint. Define the reward, the horizon, and the constraint.
Sourcing & signal
Hunting for leakage
Data & labels
Auction logs with what you bid, whether you won, and what it cost. Censored: you do not see the clearing price when you lose.
Create features
Represent
State features: budget remaining, time remaining, pacing so far.
Temporal splits & backtesting
Split
Temporal, by campaign.
Dumb baseline
Logistic regression
Gradient boosting
Calibration
Bandits
Contextual bandits for allocation across a small action set with fast feedback.
Bandits
Online learning
Reinforcement learning
Only when actions have long consequences. Train offline from logs first; online RL against a live auction is expensive.
Value-based RL
Policy-based RL
Offline & model-based RL
Model
Fixed bids or a proportional pacing rule; simple controllers are hard to beat. Bidding needs calibrated value and win-rate predictions, which are their own playbooks.
Train
Nothing unusual here.
Off-policy evaluation
Evaluate offline
Off-policy evaluation from logs before any live test.
A/B testing & interleaving
Evaluate online
A budget-split A/B.
Serving & release
Monitor & retrain
Ship & monitor
Guardrails: spend caps and fallbacks to the baseline controller. Monitor constraint violations, not just reward.
Mini Map