← All insights

"Seems Fine" Isn't a Metric: The Quality Gate to Build Before AI Workflow #2

A mid-year case for replacing vibes-based AI evaluation with a real quality gate: what evaluating on vibes looks like, the four-part gate that fixes it, and why PolicyFlow's three-month build was fast because of its eval harness, not in spite of it.

Horizon Two LabsAI THAT SHIPS"Seems Fine" Isn't a Metric:The Quality Gate to BuildBefore AI Workflow #2horizontwolabs.com

How Do You Measure AI Workflow Quality in Production?

Quick test. Your first AI workflow has been live for a few months. Did its output quality get better or worse in the last thirty days?

If your answer involves the words "I think" or "nobody's complained," you're not measuring quality. You're vibing it. And that same shrug is quietly deciding whether your next three AI projects go anywhere.

Most teams shipping their second workflow right now are carrying this exact gap. The first build got real scrutiny: everyone read the outputs, everyone had opinions, the demo got applause. Then it went live, attention moved on, and "evaluation" collapsed into a support channel that stays mostly quiet. Silence became the metric.

Silence is not a metric. Silence is what drift sounds like.

What Evaluating on Vibes Actually Looks Like

Nobody thinks they're doing this, so it's worth being specific. You're evaluating on vibes if:

  • Quality checks are somebody skimming transcripts when they remember to. That's a habit, not a gate. It stops happening the week that person gets busy.
  • "It's working" means "no one has complained this week." Users rarely report AI failures. They quietly stop trusting the output and route around it.
  • The bar for shipping a prompt change is that it looked good on the three examples you tried. Three examples is a demo. It tells you nothing about the four hundred real cases from last month.
  • Nobody can say what the workflow's failure rate was in June versus July. Not because the number is bad. Because the number doesn't exist.

None of this feels dangerous on workflow #1. It becomes dangerous the day you propose workflow #2.

Why This Is the Thing Standing Between You and Workflow #2

Half the year is behind you, and this is the season where teams decide what to scale. Here's the problem: every scaling decision is secretly an evaluation question.

Should we extend the intake workflow to a second document type? Depends how it performs on the first one. Can we cut the human review step to save cost? Depends what the error rate actually is. Did last month's model swap make things better? Depends against what baseline.

Evaluate on vibes and every one of those questions gets answered with a guess. Guess wrong in one direction and you kill a workflow that was working. Guess wrong in the other and you scale a failure mode across two more systems, which is how one mediocre pilot becomes three. We've watched that second one play out more often, because the second workflow ships on borrowed confidence while everyone is still celebrating the first.

This is the same gap that keeps AI pilots stalling before production in the first place: "it worked when we watched it" was the standard, and nobody built the thing that watches when you don't.

What a Real Quality Gate Looks Like

The fix is not a dashboard with forty charts. A working quality gate for a production AI workflow is four boring things:

A labeled eval set built from your real traffic. Fifty to a couple hundred actual cases from production, each marked with what the right answer was. Not synthetic examples. The weird, malformed, half-in-another-language cases your users actually send.

A score that runs on every change. Prompt tweak, model swap, retrieval change: it runs against the eval set first, and the number gets compared to the last number. That's the gate. If the score drops, the change waits.

A drift check on live traffic. Sample a slice of real outputs on a schedule and score them. Models change under you, input mix shifts with the season, and drift never announces itself. You catch it by looking on purpose.

A human-review rate you track like a cost. What fraction of outputs need a person to fix them? That number trending down is what ROI looks like. That number unknown is what vibes look like.

That's it. It's not research infrastructure. Most of it is a spreadsheet's worth of labeled cases and a script that refuses to let "seems fine" ship.

The Proof This Pays for Itself

When people hear the PolicyFlow numbers, one senior engineer with an AI co-pilot shipping in three months what a four-person team would have needed a year for, they assume the speed came from typing faster. It didn't. It came from being able to trust changes quickly. The build produced an eval harness early, which meant every iteration got judged by the harness in minutes instead of by a committee re-reading outputs for a week. Fast feedback is what made the fast timeline safe.

That's the part teams skip when they try to copy the speed without the gate. The co-pilot generates code and prompts quickly either way. Whether you can ship what it generates depends entirely on how fast you can tell good from bad.

Do This Before You Greenlight Workflow #2

A realistic first step, sized for one focused week: pull fifty real cases from your first workflow's production traffic. Label the right answers. Score today's system against them and write the number down. You now have a baseline, which puts you ahead of most teams currently pitching their second workflow to a budget owner.

Then make the rule: nothing ships a change without beating the baseline, and next quarter's scale-up conversation starts from the trend line, not from whoever's vibe is loudest in the room.

If you'd rather not build that muscle alone, this is exactly what our AI Enablement Retainer exists for: standing evaluation, quality gates, and governance across your AI footprint as it grows from one workflow to several, so the second build inherits a gate instead of a guess.

Get the PolicyFlow Case Study

We wrote up how the PolicyFlow build used evaluation to turn a three-month timeline into a trustworthy production system, and what the harness actually checked. Download it at horizontwolabs.com, then run the fifty-case exercise against your own workflow. If the baseline surprises you, better to find out now than three pilots from now.

Building AI into real operations is what we do.

Start a conversation

Have a workflow worth fixing?

If something in your operation takes four hours that should take two minutes, that's where we start. No pitch required.