Your Eval Dashboard Is Green. Your CFO Still Killed the Pilot. Here's Why.
Model evaluation checks whether the AI is right. It doesn't check whether the workflow is cheaper than the human process it replaced, and that gap is what kills AI pilots in budget review. This post walks through the cost-per-decision framework to fix it, using PolicyFlow's build economics as the model for run-cost discipline.
The number that kills pilots isn't on your eval dashboard
Here's the meeting nobody warns you about. Your eval harness shows 94% accuracy, hallucination rate is low, latency is fine, every test case passes. You walk into the budget review confident. Then someone from finance asks: "So what does this cost per claim compared to what we do now?" And you don't have an answer, because nobody built that ledger. The pilot doesn't die because the model is bad. It dies because you measured the wrong thing.
Model evaluation and economic evaluation are two different disciplines, and most teams only build the first one. Accuracy, pass/fail rates, and latency tell you the model works. They tell you nothing about whether replacing the human who used to make that call actually saves money. That gap is exactly where CFOs kill projects, and it's the gap that gets exposed hardest during pre-budget season, when every AI line item has to justify itself against next year's plan.
What model evaluation actually measures (and why it stops short)
A solid eval harness checks: does the model get the right answer, how often does it hallucinate, how fast does it respond, does it pass your test suite. That's necessary. It's also incomplete, because none of it accounts for what happens after the model outputs a decision.
If you want the deeper mechanics of why quality gates stop at correctness, we've written about how to measure AI workflow quality in production and why "tasks automated" is a vanity metric when three humans are still in the loop checking the work. The short version: correctness is table stakes. Cost is the actual argument.
What a real cost comparison has to include
A defensible cost comparison isn't inference spend versus a headcount line. It's:
- Human review time still required. Even a well-tuned model needs a human on edge cases, exceptions, and anything near a confidence threshold. That time has a fully-loaded cost, and it doesn't disappear just because the workflow is "automated."
- Cost of errors that slip through. Every workflow has a false-negative rate. What does it cost when one gets through, and how often does that happen at your actual volume?
- Ongoing eval and governance overhead. Production AI isn't a one-time build. It needs monitoring, re-evaluation as data drifts, and governance documentation. That's an ongoing cost, comparable to (and different in shape from) a static headcount cost, and it belongs in the same ledger.
- The old process's real cost, not its sticker cost. Most teams underprice the manual workflow they're replacing because they never accounted for rework, escalations, and the manager time spent unblocking exceptions.
If any of these categories is missing from your comparison, you don't have a cost case. You have a demo with a spreadsheet attached. For a deeper look at why workflows that look automated often aren't, see why tasks automated doesn't mean cost automated.
PolicyFlow shows the discipline, not just the headline
PolicyFlow is the case we point to for build economics: one senior engineer plus an AI co-pilot shipped in three months what would've taken a four-person team twelve months to build. That's a real, specific number, and it's a build-cost story: it tells you what it cost to ship the thing.
The mistake is stopping there. The same rigor that produced that build-cost number has to apply to the running cost of the workflow once it's live. Shipping fast and cheap doesn't automatically mean operating cheap. A workflow that took three months to build can still lose money in production if the human-in-the-loop review time wasn't priced correctly, or if the exception rate at real volume is higher than it looked in the pilot. Build cost and run cost are two different ledgers. Budget reviews care about both, but they care more about the second one, because that's the number that recurs every quarter.
A cost-per-decision ledger you can build before your budget review
Here's the framework we use with clients before they walk into finance. Build three numbers:
- Cost per decision, old process. Fully loaded: salary/hour of the person making the call, average time per decision, rework rate, escalation cost.
- Cost per decision, AI process. Inference cost per decision, plus the pro-rated cost of the humans still required to supervise and handle exceptions, plus your share of ongoing eval/governance overhead (an AI Enablement Retainer style cost, not a one-time build cost).
- Break-even volume. At what decision volume does the AI process's fixed and marginal costs cross below the old process's marginal cost? Below that volume, automation may not win yet. Above it, it should.
This is the same ledger structure we build in every AI Opportunity & Readiness Sprint, and it's the artifact that survives a CFO's second question, not just the first one.
The governance angle enterprise teams can't skip
For CTOs and CISOs, there's a second reason this matters beyond the budget meeting. NIST AI RMF and ISO 42001 don't just ask whether your model is accurate. They expect you to document the operational and cost assumptions behind an automated decision: what happens on exceptions, who supervises, what the failure cost looks like. A quality gate that only measures correctness is also a governance blind spot. If your documentation stops at accuracy, you're not audit-ready, regardless of how clean your eval scores are.
Build the ledger before finance asks for it
A green eval dashboard and a real cost case are not the same artifact, and budget season is when that difference gets expensive. If you're staring down a 2027 budget review and you don't have a cost-per-decision number for the workflow you're running, that's the gap to close first, not the eval suite.
We built a cost-per-decision worksheet for exactly this moment: it's the same tool we use in every AI Opportunity & Readiness Sprint to map a workflow's real economics in two weeks. Download it, run your own numbers, and walk into your budget review with the ledger finance actually wants to see. Start at horizontwolabs.com.
Frequently asked
What's the difference between model evaluation and economic evaluation for AI?
Model evaluation measures whether the AI gets the answer right: accuracy, hallucination rate, latency, pass/fail on test cases. Economic evaluation measures whether the workflow is actually cheaper, factoring in the fully-loaded cost of the human review, rework, and exception handling still required to run it in production. A green eval dashboard says nothing about either.
How do you calculate cost per decision for an AI workflow?
Add up the fully-loaded cost of the old process per decision (labor, review time, error rate) and compare it to the new process per decision (inference spend, remaining human supervision time, cost of errors that slip through, plus your share of ongoing eval and governance overhead). Divide by volume to find the break-even point where automation actually wins.
Does NIST AI RMF require documenting the cost of an AI decision?
NIST AI RMF and ISO 42001 both expect you to document the operational context behind an automated decision, not just its accuracy. A quality gate that only tracks correctness leaves that documentation gap open, which is a problem in an audit long before it's a problem in a budget meeting.
Building AI into real operations is what we do.
Start a conversation