← All insights

Why Your NIST AI RMF and ISO 42001 Compliance Breaks on Workflow Two

Governance frameworks like NIST AI RMF and ISO 42001 don't fail teams; teams fail to map them to specific agent decisions before building, which is why the second AI workflow trips an audit flag the first one never hit.

Why Your NIST AI RMF and ISO 42001 Compliance Breaks on Workflow Two

The governance review that passed once won't pass twice

Here's the pattern we keep seeing. A team ships its first AI workflow, something contained like summarizing intake forms or classifying support tickets. It goes through governance review, gets a green light, ships. Six weeks later they build workflow two: an agent that reads a customer record, decides on an action, and writes back to a system of record. Same team, same NIST AI RMF checklist, same ISO 42001 controls on paper. This time it trips an audit flag nobody scoped for.

Nothing about the framework changed. What changed is the blast radius. Workflow one could produce a wrong summary. Workflow two can take a wrong action. Those are different risk categories, and treating them under the same governance pass is the mistake, not the frameworks themselves.

We wrote about this exact failure pattern in Workflow One Worked. Workflow Two Is Stuck, and the root cause is almost always the same: nobody documented the decision logic granularly enough to know which parts of governance needed to scale with the agent's permissions.

What "mapping governance to agent decisions" actually means

Most teams treat NIST AI RMF and ISO 42001 as system-level compliance: you evaluate "the AI system" once, check some boxes, move on. That works for a single-purpose classifier. It falls apart the moment you have an agent making multiple distinct decisions, because each decision carries different risk.

The fix is to map governance per decision, not per system. For every discrete action an agent can take, ask:

  • Read data: What's the govern-level policy on what this agent can see? This is largely a NIST "govern" function question, i.e. is access scoped and documented before anything runs.
  • Call a tool: What could go wrong if this tool call is malformed or manipulated? This is "map" territory, identifying the risk before it happens.
  • Write to a system: How will you detect if a write was wrong, and how fast? This is "measure," and it needs actual instrumentation, not a policy document.
  • Escalate to a human: What's the fallback when the agent isn't confident? This is "manage," the operational response when something in map or measure trips.

Each of those four actions gets a named RMF function and a named ISO 42001 control before a line of orchestration code ships. Not after. If you can't say which control applies to your agent's write permission, you haven't scoped the workflow, you've just built it and hoped.

Where OWASP LLM Top 10 does the real work

NIST AI RMF and ISO 42001 are abstract by design; they're frameworks meant to apply across industries and use cases. That abstraction is exactly why teams struggle to operationalize them. The OWASP LLM Top 10 is the bridge.

Take excessive agency, OWASP's term for an agent that's been granted more capability (tool access, write permissions, autonomous decision-making) than the task actually requires. That's not a compliance checkbox, it's a specific architectural decision: did you give this agent write access to the billing system when it only needed read access to answer questions about it? Once you name it as excessive agency, you know exactly which control applies: scope permissions to the narrowest set the task requires, and document why each permission exists.

Same with prompt injection. It's not a vague "security risk," it's a specific failure point where untrusted input (a customer email, a scraped webpage, a tool's output) manipulates the agent's next action. Mapping that risk tells you which measure and manage controls you need: input validation, output monitoring, and a defined escalation path when the agent's behavior looks off. We go deeper on this specific failure mode in NIST AI RMF and ISO 42001 Won't Stop Your Agent From Deleting the Wrong Records, and on scoping data access itself in Your AI Agent's Data Boundaries Are Probably Undocumented.

Scope this during the Readiness Sprint, not the Build

This is the part that actually matters for planning: where in your build process does this mapping happen?

We do it during the AI Opportunity & Readiness Sprint, the two-week engagement before anyone writes orchestration code. At that stage, mapping every planned agent decision to its RMF function and ISO control is cheap. It's a whiteboard exercise, a document, a conversation with legal or compliance before there's a deadline attached to the answer.

Do the same mapping during the Pilot-to-Production Build, after the agent architecture is already decided, and it's a different exercise entirely. Now you're retrofitting permissions, rewriting tool integrations, and explaining to an auditor why write access was granted before anyone scoped what happens when it's wrong. That's not a flag, that's a finding. And findings cost more engineering time to remediate than the mapping would have cost to do upfront. This is the same sequencing problem we describe in our approach to shipping one workflow before scaling to the next: the order you do things in is itself a risk decision.

The line item to put in next year's AI roadmap

If you're building the case for next year's AI budget right now, governance mapping is a line item, not a footnote. It's easy to scope the cost of a Pilot-to-Production Build. It's harder to scope "the cost of finding out mid-year that workflow two needs a different risk review than workflow one," because most teams don't know that cost exists until legal asks a question they can't answer fast.

Putting governance mapping into the plan now, as a defined step in every future workflow (not a one-time audit), is the difference between a roadmap that survives contact with your second agent and one that gets rewritten in a hurry when it doesn't.

Want to see how this mapping actually gets scoped, decision by decision, against a real agent architecture? Read the full case study and download the governance mapping framework to see how we run it inside the AI Opportunity & Readiness Sprint, before build starts, not after launch.

Frequently asked

How does NIST AI RMF apply to AI agents specifically?

NIST AI RMF's four functions (govern, map, measure, manage) apply per decision an agent makes, not per system. A read-only lookup needs govern-level policy. A write action or tool call needs map (what could go wrong), measure (how you'll detect it), and manage (what happens when it does) defined before that code ships.

How does the OWASP LLM Top 10 relate to NIST AI RMF and ISO 42001?

OWASP LLM Top 10 categories like excessive agency and prompt injection are the concrete failure modes that abstract frameworks are trying to prevent. Mapping an agent action to 'excessive agency risk' tells you exactly which RMF function and ISO 42001 control apply, instead of guessing from a checklist.

Why does AI governance that passed for one workflow fail on the second?

The first workflow is usually narrow and read-only, so a general governance review passes easily. The second workflow adds tool access or write permissions, which changes the blast radius of a mistake. The framework didn't change. The thing you were supposed to scope against it did.

Building AI into real operations is what we do.

Start a conversation

Have a workflow worth fixing?

If something in your operation takes four hours that should take two minutes, that's where we start. No pitch required.