The five stages, and how to tell which one you are in
Place yourself by what your systems do, not by what your documents say. The tell for each stage is a question you either can or cannot answer today.
- Stage 1, ad hoc. People use whatever assistant they found, on whatever data is to hand. The tell: nobody can produce a list of the AI in use, and the first honest inventory surprises someone.
- Stage 2, documented. An acceptable use policy exists, along with an approval form and a register. The tell: the register is maintained by hand, it is already out of date, and nothing in the execution path knows the policy exists.
- Stage 3, observed. Traffic runs through something that records it, so prompts and responses can be reviewed after the fact. The tell: you can answer what happened yesterday, but nothing stopped it happening.
- Stage 4, enforced. Policy is applied while the work runs. A request that breaks a rule is refused before it reaches a provider, and an action an agent proposes is decided before it takes effect. The tell: you can name a refusal from last week and produce the record of it.
- Stage 5, governed at scale. New use cases inherit the controls by default rather than by project. The tell: a new agent is governed on its first day because the pipeline it runs in is not optional, and the evidence looks the same across every team.
Why organisations park at documented
The reason is usually staffing rather than intent. The IAPP's AI Governance Profession Report found that only 1.5% of surveyed organisations described their AI governance staffing as sufficient. A policy document is cheap. Putting a control in the path of a live request costs engineering time nobody budgeted, and it makes someone accountable for a refusal. NIST's AI Risk Management Framework, published in January 2023, splits the work into four functions: Govern and Map, then Measure and Manage. Stage 2 organisations have usually done the first pair on paper while the second pair stays theoretical.
Agents move the whole curve left
Earlier versions of this model, including the one that used to sit on this page, treated agent governance as a stage 5 concern. That was wrong. OWASP published a Top 10 dedicated to agentic applications on 9 December 2025, which is a reasonable marker for when agent risk stopped being a frontier topic. An agent that acts is a stage 4 problem the day it exists, because there is nothing to observe after the fact once an email has gone out or a record has changed. So if you are at documented on paper and already have an agent in production that can act in another system, score yourself at stage 1 for that agent and plan accordingly. NIST's Generative AI Profile, AI 600-1, published in July 2024, is the useful companion: it names twelve GenAI risk categories, and what you enforce at stage 4 is worth checking against that list rather than against a house definition of harm.
Moving from documented to enforced
This is a sequence, and skipping a step is how programmes end up with a control nobody trusts.
- Pick one use case with a real owner. Not the estate.
- Write down what the agent may complete alone and what stops for a person, in the owner's words rather than in policy language.
- Put the check in the path of the action, so a refusal happens before the action rather than being reported after it.
- Decide what happens when the check itself cannot answer. Anything that changes something elsewhere should stop for a person; a read-only lookup can usually proceed.
- Run the failure cases on purpose: a policy refusal, a document full of personal information, a connected system that is down, an approval nobody answers.
- Read the resulting record with the risk owner before you give the agent anything more to do.
What enforced looks like in a working system
Difinity.ai shows how stages 4 and 5 can work in practice. A turn runs in a fixed order that lives in the code rather than in a configuration option: the incoming message is checked, routing happens on the redacted text, the model is called, and the answer is checked before it is returned. A refusal at any stage stops the turn, and a request the guardrails refuse never reaches a provider. An agent run puts a loop around that same sequence rather than adding a separate late-maturity capability. The agent proposes an action, the tool gateway decides and holds the credential, and the result comes back to the agent; an action that needs a person stops the run until that person answers. Chat and agent runs share one pipeline, and whether a run is recorded is decided by the route rather than by the caller, so there is no request field that switches the trail off. That inheritance is what carries an organisation from stage 4 to stage 5. Governed run records can contribute operational evidence to wider EU AI Act, ISO/IEC 42001, risk, and audit processes. Difinity does not determine that an organisation or AI system is compliant, and it does not provide ISO/IEC 42001 certification.
Where this model is weak
Two caveats worth stating. These five stages are a synthesis informed by NIST's AI RMF and by ISO/IEC 42001:2023, whose clause structure runs from context and leadership through to improvement. They are not a named external framework, and no auditor will recognise 'stage 3' as a defined term, so use the model to argue for the next piece of work rather than as an assessment result. The second caveat is the IAPP figure quoted above. That report has appeared in more than one edition, the numbers shift between them, and at least one edition was co-published with a vendor in this market. Open the current edition before the number lands in a board pack.
Frequently asked questions
What is the difference between observed and enforced?
Observed means the activity is recorded and can be reviewed later. Enforced means a rule is applied while the request or the action is happening, so a refusal prevents the outcome instead of describing it. Logging and monitoring tooling delivers the first; ask whether it can also refuse the action.
Can we skip the observed stage?
Partly. If you are putting enforcement in the execution path anyway, you get the record as a by-product, so building a separate observation layer first is often wasted effort. What you cannot skip is knowing what AI is already running, which is the inventory work at stage 1.
How long does it take to move from documented to enforced?
For one bounded use case with an owner who can make decisions, weeks. For a whole estate, the honest answer is that it never finishes as a single project, which is why stage 5 is defined as inheritance rather than as coverage.
Does ISO/IEC 42001 map onto these stages?
Loosely. The 2023 standard is structured as a management system, running from context and leadership through planning, operation, performance evaluation and improvement, so it assumes something closer to stage 4 exists to evaluate. It does not define maturity stages, and nothing here is an audit outcome.
Sources and further reading
- NIST AI Risk Management Framework 1.0, January 2023 (opens in a new tab)
- NIST Generative AI Profile, AI 600-1, July 2024 (opens in a new tab)
- OWASP Top 10 for Agentic Applications 2026, published 9 December 2025 (opens in a new tab)
- ISO/IEC 42001:2023, AI management systems (opens in a new tab)
- IAPP, AI Governance Profession Report (check the current edition before quoting figures) (opens in a new tab)