Home/Guides/How do you evaluate an AI agent governance platform?
Guide

How do you evaluate an AI agent governance platform?

Practical criteria, tests and decision rules for evaluating an AI agent governance platform, with a 90-minute script anchored in OWASP's agentic risks.

Why agents need their own evaluation criteria

Governance built for models asks whether a system is accurate, fair to the people it affects and properly documented. Those questions still matter. An agent adds a different one: what is this thing allowed to do right now, while it is mid-run and about to call a tool? OWASP published its Top 10 for Agentic Applications on 9 December 2025, and the categories ASI01 to ASI10 are the clearest statement of what changed. Goal hijack, tool misuse, memory poisoning, rogue agents, abuse of an agent's identity or privileges: none of those are prompt-safety problems. They were compiled from real 2025 incidents, including EchoLeak, a zero-click data exfiltration issue tracked as CVE-2025-32711.

The six criteria

Use an external standard as the spine, because a vendor rubric written by the vendor is not one. Each criterion below answers a named risk category from OWASP's agentic list, or Excessive Agency from the OWASP LLM Top 10, which covers excessive functionality, excessive permissions and excessive autonomy.

  • A distinct identity and a named owner for every agent, answering identity and privilege abuse, ASI03.
  • Authority scoped to the job rather than to the integration account, answering excessive agency, LLM06.
  • A decision in the path of each consequential action, answering tool misuse.
  • A credential the agent cannot reach, answering rogue agents, ASI10.
  • A person in the loop above a threshold you set, answering human-agent trust exploitation, ASI09.
  • An append-only record of what was attempted, what was refused and what resulted, which is where goal hijack and memory poisoning surface before anywhere else.

The test for each one

Criteria are cheap. These are the questions that produce different answers from different products.

  • Identity: ask which record in your identity system represents the agent, and who gets called when it misbehaves.
  • Authority: attempt an action the job does not need, then check whether anything downstream changed.
  • Decision in the path: ask for a demonstration of a refusal rather than a success.
  • Credential: ask what the agent process holds while a tool call is in flight.
  • Human approval: check who may approve, and whether an approval can be replayed on a later action.
  • Evidence: hand one run to a reviewer who did not build the agent and ask them what happened.

Identity and least privilege, stated properly

Least privilege for agents is not a smaller API key. The useful formulation is that an agent's authority is the intersection of what its agent version binds, what the person or application it acts for is entitled to, and what the business context permits. An intersection, never a union. If a platform computes it as a union anywhere, an agent inherits the widest permission in the chain and the whole model collapses quietly.

Two follow-ups worth asking

What happens when a tool server runs under its own identity rather than the person's, which can hand a caller reach that person does not have? And what happens to an agent's authority when the person it acts for changes role or leaves the organisation? Both are ordinary situations, and both expose whether authority is computed at the moment of use or copied once at setup.

The decision has to sit in the path

This is the criterion most often claimed and least often met. Reading a trace after an agent acted, evaluating it continuously, alerting on a threshold: all useful, all after the fact. Read every runtime page the same way and sort the verbs. Reading a trace, evaluating it and raising an alert are one class of thing. Refusing an action while the run waits is another. A product can be excellent at the first and have no capability at the second, and the marketing word for both is often the same.

Evidence a reviewer can read on its own

An engineering trace and an accountable record are different artefacts with different audiences. Security wants the attempted violations and risk wants to see which control applied, while the business owner only cares about the outcome for a named customer. Require that the record cannot be edited, that a repeated write is rejected rather than stored twice, and that refused actions are recorded as carefully as completed ones. A denied refund or a blocked outbound call is usually the most interesting line in the whole run.

A 90-minute evaluation you can run yourself

Give every shortlisted product the same small workflow and the same evidence request. One agent, one bound customer or case, one permitted read, one permitted write, one protected field, one action outside the agent's authority, one dependency you take away mid-run.

  • First 15 minutes: register the agent with an owner, a job, the systems it touches and the actions it may take.
  • 15 to 30: run the permitted path, then verify the result in the target system rather than on the vendor's screen.
  • 30 to 45: put realistic sensitive data through it and inspect what each destination received.
  • 45 to 60: attempt the action just outside the agent's authority and confirm nothing changed downstream.
  • 60 to 75: take away a dependency mid-run, such as the policy service or the target system, and watch the recovery path.
  • 75 to 90: hand the run record to a reviewer who configured none of it and ask them what happened.

Decision rules for reading the answers

The evaluation only helps if you know what each answer means. These four rules settle most of it.

  • If the product reads traces and cannot refuse an action, treat it as evidence tooling and keep looking for the control.
  • If the agent holds the credential for the target system, every policy the product applies is advisory, whatever the demo showed.
  • If governance has to be configured by hand for each new agent, price the tenth and the hundredth agent, not the first.
  • If a reviewer needs the person who built the agent to explain a run, the evidence model is incomplete regardless of how much it logs.

How Difinity.ai answers its own criteria

It is fair to ask a vendor publishing a criteria list to answer it, so here it is concretely. The decision sits at the tool gateway, which holds the credential, decides before an action runs, and is not reachable from the internet. No customer calls it and neither does an agent. Authority is structural rather than a per-feature setting: it is the intersection of what the agent version binds, what the caller is entitled to, and what the use case permits, and an agent version can bind at most 32 tools. Under each bound tool, the author writes when the agent may use it in their own words, and a judge reads that policy on every effectful action.

What Difinity records, and who may approve

Evidence is append-only and cannot be edited. The run trail records each guardrail verdict, each proposed action, the gateway's decision, approvals asked for and answered, and the outcome, and a repeated write is rejected rather than stored twice. Only the person the agent is acting for may approve one of its actions, an approval is single use, and a run nobody answers expires. Governed run records can contribute operational evidence to wider EU AI Act, ISO/IEC 42001, risk and audit processes. Difinity does not determine that an organisation or AI system is compliant, and it does not provide ISO/IEC 42001 certification.

What these criteria leave out

The programme-level work. Registering AI systems, classifying risk, mapping controls to a framework and keeping assessments current are real obligations, and a runtime control does not do them. Most enterprises will end up running something for the programme and something for the runtime, which is a sensible outcome rather than duplication.

When these criteria change

OWASP's agentic categories were compiled from 2025 incidents and will be revised as the attack pattern moves. Several governance products also have enforcement work on their roadmaps, so a product that only reads traces today may decide tomorrow. Re-run the six tests at renewal rather than trusting the answer you got at purchase. One note on standards, since it comes up: there is talk of an agent-specific profile extending the NIST framework, and we could not confirm a published document on NIST's own site, so this page cites the AI Risk Management Framework 1.0 from January 2023 and the Generative AI Profile from July 2024, both of which are published.

Frequently asked questions

What should an AI agent governance platform do?

Give each agent an identity and an owner, bind it to an approved scope of action, decide whether a proposed action may run, keep credentials away from the agent, and record what was attempted and what resulted.

How is this different from AI governance?

AI governance covers the programme: inventory, risk classification, policy, assessment work and accountability. Agent governance covers the operating layer underneath it, where an agent is calling tools and changing records in real systems.

Does an agent need approval for every action?

No, and requiring it defeats the purpose. Set a threshold. Low-impact actions can run within their configured authority, and actions above the line stop for a person who can see the full action and its arguments before deciding.

Which standards should an evaluation cite?

The OWASP Top 10 for Agentic Applications from December 2025 for agent-specific risk, the OWASP LLM Top 10 for excessive agency, the NIST AI Risk Management Framework 1.0 with its Generative AI Profile for programme structure, and ISO/IEC 42001:2023 if you run an AI management system.

How long should an evaluation take?

Ninety minutes per product, using one real workflow with a permitted action and an out-of-authority action. Anything longer usually means you are being shown features rather than testing a control.

Sources and further reading

Have an agent that needs production authority?