Classification confidence

How EngLedger scores its confidence in each classification, what the high/medium/low tiers and the grey zone mean, and why the number matters for a defensible ledger.

Updated 18 July 2026

Every classified number in EngLedger carries a confidence: not a guess about whether the classification is right, but a record of how it was decided. A number a person set or a rule matched is trustworthy; a number that fell through to a default is not, and the ledger says so rather than presenting both at equal strength. This is the honest-numbers principle applied to classification.

How confidence is scored

Confidence comes straight from the classification source — the same source the classification engine records on every entry. It isn't a separate calculation you can tune:

How it was classifiedConfidence
Manual decision (someone set it)1.0
Matched an explicit rule0.9
Fell back to your default0.4
Unclassified / provider-typed only0.1

A manual override is the most trustworthy because a person made the call; a rule match is nearly as good because someone wrote a rule that explains it. A fallback means nothing you configured matched, and an unclassified entry has even that much less behind it.

Tiers and the grey zone

Scores roll up into three tiers:

TierScoreIn practice
High≥ 0.8Manual and rule-matched work
Medium0.5 – 0.8
Low< 0.5Fallback and unclassified work

Everything below 0.5 is the grey zone: effort that reached the ledger without a rule or a person behind it. It's not necessarily misclassified — it's unexplained, and unexplained money is what an auditor asks about. The grey zone is the review queue: the shorter it is, the more of your capitalisable/expensed split rests on decisions someone can defend.

The Confidence Risk report

The Confidence Risk report puts a dollar figure on the grey zone:

  • High confidence — the share of classified cost that's rule- or manual-backed. Aim for 80% or above; below that, your rules have gaps.
  • Grey-zone cost — the money classified by fallback or unknown source. This is the figure to drive down before you rely on the split.
  • Review queue — the count of grey-zone entries waiting for a human pass.
  • Ledger coverage — the classified entries and cost the report is built from.

The Grey Zone review table lists each low-confidence entry — project, ticket, classification, source and cost — so you can work it top down: extend a rule to cover a recurring pattern, or set the classification manually. Excluded ticket types are left out entirely rather than counted as low confidence.

Why it matters

Confidence is the trust layer under every other number. A month can show a perfectly reasonable capitalisation rate that's actually built on fallbacks — the confidence score is what stops that passing unnoticed. It's why period close won't read as ready while low-confidence entries remain, and why the R&D capitalisation workflow treats driving the grey zone down as a step in its own right. A smaller number you can walk backwards beats a tidier one you can't.