Field notes on building AI that ships.
Practical lessons from designing and deploying AI into real operations — the unglamorous engineering that turns a demo into a production system.
NIST AI RMF and ISO 42001 Won't Stop Your Agent From Deleting the Wrong Records
NIST AI RMF and ISO 42001 give you governance vocabulary and an audit trail, but they don't scope what your agent can actually delete, merge, or bulk-edit in production; this post explains why a written authorization map, not framework compliance, is what actually prevents excessive agency incidents, and why mid-year (when teams scale from one workflow to several) is exactly when that map needs to exist.
Read article →Your AI Workflow Works. Is It Actually Cheaper Than What It Replaced?
Most evaluation harnesses measure whether an AI workflow works, but not whether it's cheaper per transaction than the manual process it replaced; this piece walks through how to build a cost-aware eval gate before scaling one workflow into many.
Read article →"Seems Fine" Isn't a Metric: The Quality Gate to Build Before AI Workflow #2
A mid-year case for replacing vibes-based AI evaluation with a real quality gate: what evaluating on vibes looks like, the four-part gate that fixes it, and why PolicyFlow's three-month build was fast because of its eval harness, not in spite of it.
Read article →Excessive Agency Is the Security Failure Nobody Scoped: How to Map What Your Agents Can Actually Do
Excessive agency, the OWASP LLM Top 10 risk of AI agents holding more permission than their task requires, is a scoping failure that shows up hardest when teams scale a working pilot into new workflows. This post gives a concrete audit framework (enumerate access, ask the worst-case question, require human approval on irreversible actions) and ties it to how Horizon Two Labs scopes access from day one.
Read article →Why Some Teams Are Shipping Their Third AI Workflow While Others Are Still Stuck on Their First
A mid-year, no-hedging look at why some teams are already on their third production AI workflow while others are stuck rebuilding from zero on their second, using PolicyFlow's three-month build as proof that reusable eval, security, and integration patterns (not better models) are what compound.
Read article →Why Your Second AI Workflow Is More Dangerous Than Your First
The real mid-year AI decision isn't which workflow to scale next, it's whether your first production workflow is a system or just one engineer's tribal knowledge, and this piece makes the case for evaluation harnesses (and the PolicyFlow proof point) as the fix.
Read article →How to Get an AI Pilot Into Production Without Getting Stuck in the PoC Trap
A direct guide to escaping the AI proof-of-concept trap, using PolicyFlow and SoloStream as concrete proof that narrow scoping, evaluation harnesses, and right-sized governance are what actually get a pilot into production.
Read article →Why Your Second AI Workflow Needs Governance the First One Never Had
Your first AI workflow shipped safely because a human watched every output by hand, not because governance was working. This post breaks down why that approach collapses on workflow #2 and what a real evaluation and governance layer (harness, data boundaries, named owners) actually looks like, using PolicyFlow as proof that guardrails baked in from day one are what make speed possible, not what slows it down.
Read article →Shipping Workflow One Doesn't Teach You Workflow Two
Shipping one profitable AI workflow proves you can ship that workflow, not that you have a repeatable system. This piece breaks down why the mid-year urge to scale to three agents at once is more expensive than shipping zero, and what actually transfers between workflow one and workflow two: the evaluation harness, not the code.
Read article →Your First AI Workflow Shipped. Your Second One Is Drowning in Review Cycles. Here's Why.
Your first production AI workflow succeeded because it was small enough to skip every hard organizational question. Your second one is stalling because those questions are now unavoidable, and the fix is instrumented feedback loops and a pre-agreed quality gate document, not a rebuilt governance framework. This post explains the specific failure modes (evaluation limbo, missing eval harnesses, security gaps that compound across workflows) and what to do about them before workflow two enters review.
Read article →Production AI Monitoring for Q2: Catching Silent Failures, Cost Overruns, and Quality Drift Before They Compound
You shipped the pilot. The demo worked. The stakeholders approved it in Q1, and now it's Q2 and the workflow is live in production. Congratulations. Now the real work starts. Here's what nobody tells…
Read article →Why AI Pilots Stall Before Production
The gap between a working demo and a system your team trusts on a Monday morning is where most AI projects quietly die. Here's how we close it.
Read article →Ship One Workflow, Not a Transformation
You don't need to “transform your organization with AI.” You need one four-hour task to take two minutes instead. Start there.
Read article →