Move the numbers to match your own setup. The two budgets that matter are inference (tokens burned thinking) and spending (things the agent buys). A cap on one does nothing for the other.
| Cap | Where it must be enforced | What it protects |
|---|---|---|
| Inference ceiling | The provider console | The wall that stops a runaway loop from becoming an incident |
| Per-task estimate | Stated by the agent before the run | Surprises — a number in the transcript beats a trend on a dashboard |
| Velocity alert | The harness that dispatches work | Rate, not totals: a cap that trips on the 30th is a month of damage |
| Per-transaction cap | The card or wallet credential | One bad purchase — small enough to be annoying, not painful |
| Approval gate | Your own phone, above a threshold | Everything irreversible, and it makes the other four survivable |
Your inputs, your arithmetic — nothing here is a benchmark. Set the inference ceiling at roughly four times a measured week plus 50% headroom, and write one rule into the agent: on reaching a limit, stop and report — never downgrade or improvise to finish.