The AI governance report suite
A reference for each AI governance report: what it shows, the assumption behind every estimate, and how to read the value, spend, adoption and reliability numbers honestly.
Updated 18 July 2026
This is a tour of the AI governance report suite: what each report shows and, where a number rests on an assumption, what that assumption is. It's the reference behind the governance review loop; the governance model covers the posture, and AI spend attribution covers how spend attaches to work.
Value and ROI
ROI / Value estimates what AI usage was worth against what it cost. The value side is modelled: sessions × assumed hours saved per session × the person's rate and overhead, divided by spend to give a return. The hours-per- session figure is an assumption you set, so treat the ROI as a scenario, not a measurement — the report shows the break-even hours-per-session so you can judge whether the assumption is plausible.
Alongside it sits a measured comparison: the cycle time of AI-touched tickets against the rest. This is shown only when both groups have at least five closed tickets in the period; below that the sample is too small and the comparison is withheld. Even when shown, it's an observed correlation, not a proven speed-up — people choose when to reach for AI, and the report refuses to launder that into causation.
Spend efficiency
Cache ROI reports cache savings and hit rate, with an annual projection that is simply this month's saving × 12. It's a naïve run-rate, not a forecast: a single heavy or light month scales straight through, so read it as "at this month's rate" rather than a prediction.
Model Efficiency shows the mix of higher-cost versus standard models and an estimated saving from shifting the expensive work down — calculated as high-cost spend × 0.8, i.e. assuming most of it could move to a cheaper equivalent at roughly a fifth of the cost. Which models count as high-cost is a configurable list your organisation controls. The saving is a hypothesis to test against real workloads, not a guaranteed cut.
AI Tax expresses AI spend as a percentage of engineering cost, trended month over month with the month-on-month change. It needs engineering cost configured (hourly rates in place); without them it reads as unavailable rather than inventing a denominator. The concept is covered in full under attribution and the AI Tax.
Spend vs Throughput plots each person's AI spend against their delivery output (merged pull requests and commits) and reports the correlation between them. Correlation is not attribution: a positive figure says spend and output moved together, not that one caused the other.
Tag Spend breaks spend down by metadata tag. It needs tagged requests to show anything — with no tags present it reports that the data isn't available rather than an empty chart.
Adoption and people
Adoption shows how many people are actively using AI this period against your total headcount, with an onboarding and coverage view. "Active" means a person with metered usage in the period.
Person Usage and Teams rank spend and usage per person and per team. Per-person detail is an opt-in, coaching-framed surface by design — an estimate against an org-wide assumption, not a performance measure. The governance model explains the privacy stance, and people can always see their own usage.
MCP Adoption ranks which MCP servers are being used and trends their uptake, with the share of sessions that use any MCP server at all. It needs sessions that actually record MCP usage to populate.
Reliability
Health reports error rates and p95 latency by model. The error-rate read is banded: at or under one percent is healthy, up to three percent is worth monitoring, and above three percent is elevated — spend burned on failed requests and retries is spend with nothing to show for it.
Session intent
Session Intent shows the mix of what AI sessions were for — feature, bugfix, spike, chore, investigation or other — as a share of classified spend. Where a session links to a typed ticket the intent comes straight from that type; where it doesn't, the session is classified automatically and carries a confidence score. Intent classified with confidence below 70% is flagged as indicative rather than shown at full strength, so a shaky label doesn't get read as a firm one. The mechanism is covered under session intent.