AI spend budgets, limits and enforcement
Set an organisation AI budget, add team and person spend rules, preview and shadow-test them, and see what happens at the cap: block, restrict models, rate limit or alert — with break-glass overrides when someone needs access now.
Updated 18 July 2026
Metering tells you where AI spend went; spend control decides what happens next. On the AI Spend Control screen an admin sets budgets and rules (organisation-wide, per team or per person) and chooses what enforcement means when a cap is reached: block access, restrict to cheaper models, or alert and carry on.
The organisation budget
The root of spend control is the organisation budget: a single amount per period that applies to everyone. You set the amount in your organisation's currency, choose a period (daily, weekly or monthly) and pick when it resets: a reset hour for daily budgets, a reset weekday for weekly, a day of the month for monthly (months shorter than the chosen day reset on their last day). Period boundaries follow your organisation's timezone.
Rules for teams and people
Below the organisation budget sit rules. A rule's scope comes from its conditions: add a Team condition and it applies to that team, an Person condition and it applies to that person; rules can also condition on Model, Provider, Day of week and Time of day.
Each rule combines a limit with an action:
| Action | At the cap |
|---|---|
| Block | AI access pauses until the period resets or the budget is raised |
| Restrict models | Higher-cost models are denied; cheaper ones keep working |
| Rate limit | Requests and tokens per minute are capped: access slows rather than stops |
| Alert only | Nothing is enforced: admins are notified as spend climbs |
| Exempt | The rule's scope is excluded from broader rules |
Limits are an amount in your currency, a percentage of the organisation budget, or none at all: a rule with conditions but no limit acts as a standing policy (for example, restricting a model or rate-limiting outside working hours). Rate-limit rules set the requests-per-minute and tokens-per-minute caps directly; where several apply to one person, the lowest wins. Rules can be paused without deleting them.
The model catalogue
Rules that restrict or allow models pick from your organisation's model catalogue: the set of models your connected AI providers currently offer, with their provider and a readable name. The catalogue refreshes on its own, so new models appear and retired ones drop off without anyone maintaining a list by hand. When a restrict-models rule denies the higher-cost models, it's denying entries from this catalogue and leaving the rest available.
Preview and shadow mode
Two safeguards let you see what a rule will do before it does it to anyone.
Preview runs a rule set against current spend without saving or enforcing anything. It shows exactly who it would affect right now: which scopes would block or restrict, which people would lose which models, and the reason for each. Because a preview reflects spend at the moment it ran, editing a rule invalidates it — run it again after any change.
Shadow mode is for a rule you've saved and want to watch live. A shadow rule is active and evaluated every cycle, but it only reports what it would do — "would block", "would restrict" — without pausing or restricting anyone. It's the safe way to run a real rule against real spend for a period and confirm it behaves before letting it enforce. The live status shows shadow breaches as information, not alarms.
Rule sets are versioned. Every edit creates a new version rather than changing the active one; activating a version takes effect immediately, and you can roll back to an earlier version if a change lands wrong. Only one version is active at a time.
What happens at the cap
Enforcement fires when a scope's metered spend reaches its limit (100%, not a configurable warning percentage). Before that point, alert-only rules notify admins at 50%, 80% and 100% of the limit; enforcing rules notify at 100%. Each alert fires once per period. The notification is an email to your organisation's admin: the subject names the budget and the percentage it has reached, and the body shows current spend against the cap with a link to the spend report.
When a block or restriction lands, the affected person receives a short, neutral notification that their AI access is paused or limited and when the period resets; admins see which rule fired and the spend against the cap. Enforcement is applied on a short evaluation cycle, so a breach takes effect within about five minutes rather than mid-request.
Once a spend cap has triggered, its state holds until the period resets or an admin raises the budget: a momentary dip back under the line doesn't flap access on and off. Access-window restrictions clear at the window boundary instead.
When several rules cover the same person
Caps are ceilings, so every rule covering someone applies at once — there is no first-match winner. The outcome is the most restrictive combination:
- any rule that blocks wins, whatever the others allow;
- model restrictions intersect — an organisation-wide deny of one model and a team allow-only list leaves the allow-only list minus that model, a result neither rule produces alone;
- rate limits take the lowest value in play;
- an exempt rule or an active break-glass override clears the person entirely, whatever else covers them.
This is the opposite of classification rules, where priority decides a single winner. Priority here only picks the organisation budget when more than one candidate exists, so reordering rules does not change who is capped.
If an outcome surprises you, the dry-run preview names every scope covering a person and the reason each one fired.
What the person sees
A blocked request is refused before it reaches the model provider, so nothing is spent on it. The caller gets a refusal naming the organisation's spend controls, the reason in plain terms — the budget for the period has been reached, the request is outside allowed hours, or no model is permitted — and the date access returns.
The refusal deliberately carries no amounts. Whether people see the organisation's budget figures is your decision, not something enforcement leaks by default; the numbers stay on the admin screens and in the reports.
A restricted person is not refused at all: requests for permitted models carry on working, and only the restricted ones are declined, naming the model.
Break-glass overrides
Sometimes someone needs access now — an incident, a release, a deadline — and waiting for the period to reset isn't an option. A break-glass override grants one person full access for a fixed window. You set an expiry and a reason, and while it's active the override takes precedence over every budget and rule affecting that person.
An override lifts on its own when its window ends; you can also revoke it early. Every override and revocation is written to the admin audit log with who granted it, for whom, why, and until when — so a deliberate exception stays visible and accountable rather than becoming a quiet hole in the budget.
Reading the live status
The screen's status strip shows each scope with its spend against its limit and a state:
| State | Meaning |
|---|---|
| OK | Under the limit |
| Alerting | An alert-only rule has crossed a notification threshold |
| Restricted | Model restrictions are active for this scope |
| Blocked | AI access is paused for this scope |
A Currently affected list names the people under an active block, restriction or rate limit, with the rule responsible, any denied models and the per-minute caps in force. Rate-limited people appear under the Restricted state. The dashboard also raises an alert when anyone's access is paused or limited, linking to the spend report view behind it.
When enforcement doesn't apply
Enforcement decisions have to reach the point where AI requests are metered before they take effect, and EngLedger only pushes a change when something actually changes. If a change repeatedly fails to land, that persistent mismatch raises an admin alert rather than failing silently — a control that isn't sticking gets noticed instead of quietly doing nothing.
The bias is toward not cutting people off by mistake: if evaluation itself can't run, access is left as it is rather than blocked, with a hard budget ceiling still in place as the final backstop. Spend control should fail open, not trap someone behind a broken rule.
Rolling controls out
Sensible sequence: start with the organisation budget and alert-only rules, watch a full period, then tighten. Alert-only gives you the notification trail without anyone losing access while budgets are still guesses; once the numbers settle, convert the rules that matter to restrict or block. Rolling out AI spend visibility covers the communication side, and the honest-numbers framing on the AI Governance overview is worth holding onto: enforcement is a budget decision, not a judgement about how anyone works.