Documentation

Difinity Platform Guide

Technical guide to Hub, the Platform API, the Flow runtime, policy controls, and run evidence.

Documentation

Difinity.ai governs what AI agents and AI chat are allowed to do inside an organisation. This page describes the platform as it works today.

It is written for the people who decide whether to turn it on, and for whoever will own the first deployment. Everything below describes behaviour that is built, not planned.


What Difinity is

Difinity runs two kinds of work under one set of controls. If your agents only answer questions, most of what follows is more than the job needs. It earns its keep once an agent starts acting in systems that matter.

Governed chat is a person asking a model something. A governed agent is an assistant that can also act in other systems, such as sending an email or posting a message. Both run through the same enforcement pipeline. What differs is where the history comes from and whose credentials are used.

An organisation configures its use cases, agents and connections in Hub. People do the work in the chat and agent workspace at chat.difinity.ai. An application can also call the runtime directly.

Three things hold throughout. An agent does not set its own authority: every action is bounded by what its version binds, what the caller is entitled to and what the use case permits. Every action it takes leaves through a gateway that holds the credential and decides. What happened is written to an append-only record as it happens.

The parts and what each does

Difinity has four parts. Three are things an organisation configures or calls. The fourth is where actions actually happen, and nobody outside the platform calls it.

Hub

Hub is the administrative interface for an organisation: use cases, agents and their versions, connectors, tool servers, people, roles and permissions, billing, and the review of evidence.

Hub is an administrative interface, not the system that owns the state. It reads and writes through the Platform API.

Platform API

The Platform API at platform.difinity.ai is the system of record. It holds agents and agent versions, use cases, conversations and messages, connector and tool server registrations, approvals, the run trail, credits and metering, and identity records.

It also issues the short-lived tokens an application uses to reach the runtime.

The runtime reads what a run needs from the Platform API when the run starts, and cannot serve while the Platform API is unavailable.

Flow

Flow at api.difinity.ai is the runtime. It runs the guardrail pipeline, the model call, streaming and the agent loop. It holds no state of its own.

The tool gateway

Every action an agent takes leaves through the tool gateway (internally, Marshal). The gateway holds the credential, applies the rules the organisation set, decides, and acts.

The gateway is not reachable from the internet. No customer calls it, and neither does an agent.

An agent cannot hold a credential or reach a system directly. Every action is proposed to a gateway that decides, holds the credential and acts.

People reach all of this through the chat and agent workspace at chat.difinity.ai, and applications reach it through the API.

How a governed run works

A turn runs in a fixed order, and the order is part of the code rather than a configuration option.

  1. The incoming message is checked first. Where the use case is configured to detect personal information, detected values are replaced. The message is also evaluated against the use case's compliance rules, topic scope and blocked words.
  2. Routing happens next, on the redacted text. Where the use case selects a model automatically, the router reads the redacted message and never the original.
  3. The model is called.
  4. The answer is checked before it is returned.

A refusal at any stage stops the turn. A request the guardrails refuse never reaches a provider and is not charged.

An agent run puts a loop around that sequence. The agent thinks, proposes an action, the gateway decides, and the result comes back to the agent. The run is bounded, and an action that needs a person stops the run until that person answers.

Chat and agent runs share one pipeline. Whether a run is recorded is decided by the route, not by the caller. There is no request field that switches the trail off.

Agents

What an agent is

An agent is a named assistant that lives inside one use case and inherits that use case's policy. Several agents can share a use case. The agent itself has a stable identity and carries no configuration. Its configuration belongs to an agent version.

Agents are built in Hub. People run them in the workspace at chat.difinity.ai.

Purpose and instructions

Two fields do different jobs, and the difference matters.

Instructions are what the agent is told. They are sent to the model.

Purpose is what the agent exists to do, written by its author. It is not sent to the model. It is one of the things the gateway judges a proposed action against, so an agent built to answer questions about returns can be held to that when it proposes something else.

Versions and review

An agent version is an immutable definition: instructions, model, reasoning tier, tool bindings, knowledge and run limits. It is the unit that is reviewed, approved and published.

A version moves through Draft, In review, Approved, Published, then Superseded or Unpublished. A change note is required to submit one.

Whether a version needs review is a setting on the use case. If review is not required, submitting the version approves it directly. Who may approve is the Agents:approve permission, not a named approver group.

A conversation pins the version that is published at the moment it starts, so publishing a new version does not change a thread that is already running.

Tools, connectors and tool servers

An agent's tools are bound to the version, and each binding names the tool and the catalogue it was approved against.

Gmail and Slack are the connectors Difinity brokers. Anything else is an MCP server the organisation registers. Hub labels these Tool servers.

An administrator turns a connector on and chooses which of its tools agents may call at all. That is the organisation's ceiling. A person then connects their own account, and an agent acting for that person uses that person's credential, never anything wider than the ceiling.

An agent version can bind at most 32 tools, and every one is a deliberate choice rather than a default.

A published version keeps the tools it was approved with. Re-reading a tool server's list does not change a published agent. Picking up a new tool is a draft edit that goes back through review.

Authority is the intersection of three things: what the version binds, what the caller is entitled to, and what the use case permits. It is never a union.

Where a tool server runs under its own identity rather than the person's, Hub marks it as such, because a server like that can give a caller reach that person does not have.

Per-tool policy

Under each bound tool, the author writes when the agent may use it, in their own words. "Only email the customer on this order" is a policy.

That text is read by the judge and is never matched by a parser. Left blank, the judge has no rule of the author's to apply to that tool. That does not switch judging off for anything effectful, because the floor described below applies whatever the written policy says. A read-only tool with no written rule and no configured posture is not judged.

The gateway refuses any tool it holds no schema for. Arguments it cannot check are arguments it cannot constrain.

The judge

Every effectful action, meaning anything that changes something elsewhere, is read by the judge before it runs. A read-only tool is judged when its author wrote a rule for it or the organisation set a posture for it, and not otherwise. Where the judge runs, a first model tier on AWS Bedrock reads the call, a second tier on Bedrock reads the calls the first would not commit on, and a call neither tier can settle goes to a person.

The judge can only refuse or escalate. It can never permit something the deterministic rules refused, so its worst failure leaves the deterministic answer standing.

It never sees the conversation, because a judge given the conversation is given whatever was hidden in it.

Anything that changes something elsewhere is reviewed whatever the written policy says. Tools are marked as effectful or read-only, and the effectful ones carry a floor a policy cannot lower.

When the judge cannot decide, including when Bedrock cannot be reached, the outcome follows the tool's configured posture. An effectful tool escalates to a person whatever that posture says. A read-only tool set to permit is permitted, and a read-only tool that has a written rule but no posture escalates.

Approvals

Only the person the agent is acting for may approve one of its actions. There is no approver group, no shared queue and no delegation. A run nobody answers expires.

An approval is single use.

An "always allow" answer is keyed to the policy clause that raised the question, and it is read back the next time that clause fires. A refusal is never standing, because "no, and never again" is a policy change rather than a side effect of an approval.

The person is shown the whole action and its arguments, not a summary. The answer is signed, and the gateway checks the signature and the digest against the action it stored before anything runs.

Reasoning tier and run limits

A version carries a reasoning tier, Quick or Deep. Deep is given a larger step budget, more time and a higher ceiling per run.

The tier belongs to the approved version and not to the request, so a caller cannot turn a quick assistant into an expensive one. There is no separate step budget for an author to set.

A run that reaches its ceiling stops and reports that it did.

Chat and agent workspace

chat.difinity.ai is the chat and agent workspace. It uses the same account as Hub, and one sign-in covers both. There is no self-serve signup on it.

Chat and agents

There are two workspaces. Plain chat runs against a chat use case the organisation has configured, and a person can pin one of that use case's models or let the router choose. An agent thread is bound to one agent, whose published version decides the model, so there is no picker.

Only agents with a published version are offered.

Files and document review

A person can attach a file. Accepted types are pdf, docx, pptx, xls, xlsx, csv, txt, md, html, json and xml. Legacy .doc and .ppt files are refused.

An attachment is at most 25 MB and is kept for 28 days.

Where an attached document carries personal information, the turn pauses and the person is shown the redacted copy of each document. They can correct it. Only what they confirm is sent, and the confirmed copy stays readable in the thread afterwards.

Watching a run

The workspace shows the run in the order it happened: the checks, the model's own reasoning, each tool call, and what a tool returned. The model that answered is shown on the answer.

A turn can be stopped. Stopping ends the run at its next step rather than hiding it.

Answering an approval in the thread

An action that needs approval appears in the thread at the point the run reached it, with the tool named and its arguments in full.

The arguments shown are the restored, real ones. That is what was signed and what will run.

What the workspace does not do

A conversation can be archived, which hides it while its messages remain. There is no rename. An organisation can use the self-service export for its data, and a person can download back a file they attached, one file at a time until the file expires.

Integrating an application by API

An application can call the runtime directly instead of working through the workspace. The run is governed in the same way.

  1. An administrator registers the application in Hub and issues an API token. The token secret is shown once, when it is created, and never again.
  2. The administrator assigns the use cases that application may run against.
  3. The application exchanges its API token at POST /api/v1/auth/exchange on platform.difinity.ai and receives an identity token valid for 15 minutes.
  4. It calls Flow at api.difinity.ai with that token, naming the use case on every request as the use_case_id query parameter or the X-Use-Case-ID header.
  5. An administrator reads the evidence in Hub. The calling application does not read the trail.

The default rate limit is 60 requests per 60 seconds per application. A 429 carries X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset and Retry-After.

Request and response shapes, the error envelopes, streaming and the provider-compatible endpoints are in the API reference.

Personal information

The boundary is one sentence. The model works on redacted text, and real values are restored in the last step before an action leaves through the gateway.

Restoring real values

An agent proposes actions in redacted form, because redacted text is all it has seen. The arguments are restored immediately before the action is submitted to the gateway, because every argument rule the gateway holds is written about real values.

An action waiting for approval shows the restored form, which is what was signed and what will run.

A run that stops for an approval keeps its redaction map, so results that arrive after the answer do not reach the model with real values in them.

Checking what comes back

A tool result is checked before the model reads it. Values already hidden in that session are hidden the same way in the result, so the model sees one stand-in per value. A result that cannot be checked is withheld.

What is configured, and what this does not do

What counts as personal information is configured per use case. Defaults are seeded where they can be read, narrowed or removed, and whatever stands there is what runs. An empty list stops the detector being called.

Detection is a model, and it is not perfect. Where a use case is set to protect, a failed check refuses the request rather than passing text on. Where it is set only to report, it never refuses, because someone who asked to be told about personal information has not asked for the conversation to stop.

Where detection and replacement are enabled, the trail records the redacted stand-ins used during the run, not the original detected personal values. This applies to recorded messages, model output, tool-call arguments and results, and policy reasons. If an organisation turns personal-information detection off, the trail stores the raw text and labels the run accordingly.

Evidence

There are two separate records. They exist for different purposes and must not be treated as interchangeable. This is where most of the confusion starts.

The run trail

The run trail, which Hub calls the AI Trail, is append-only evidence of one run: the message, each guardrail verdict, each model turn, each proposed action, the gateway's decision, approvals asked for and answered, each executed action, redactions, and the outcome.

It cannot be edited. The application role holds no update grant, the table carries no policy that would allow one, and a repeated write is rejected rather than stored twice.

The actor records who vouched for the person, not merely who was named: the subject, the issuer, and whether the identity was verified. Where an application asserts an end user of its own, the trail shows that as an unverified claim by that application.

A stated reason is displayed as a claim and never as a fact. The model writes it, and under a prompt injection that means an attacker writes it.

The Configuration Log

The Configuration Log holds the activity log, which is who changed what, and the access log, which is one request against the API surface. Both are append-only, and there is no route that edits either.

A transcript is not the trail

A transcript is the person's own data and keeps the real values for their own conversation history. The trail is evidence. Neither is derived from the other and both are written. Deriving one from the other would force a choice between refusing an erasure request and destroying evidence.

What the trail does not record

Tool results are not stored by default. What is kept is the decision, the arguments, the outcome, the shape of the result and a keyed digest of it. Storing results in full is a choice an organisation makes.

Redaction entries record the kind of value and its stand-in, not the original detected personal value.

Retention and erasure

Each organisation sets its own periods. The run trail defaults to seven years and cannot be set below six months. Transcript retention defaults to while the organisation remains subscribed and can instead be set to a specified number of days. If the organisation does not choose a period, those defaults apply. When a period expires, the applicable data is deleted. Mapping rows expire with the trail records they belong to.

A legal hold pauses scheduled expiry and erasure until the hold is lifted. Chat attachments expire after 28 days. Runs waiting for approval expire after seven days. Database backups are kept for seven days with point-in-time recovery.

An organisation can use the self-service export while subscribed. The organisation enters the Terminated state only when authorised Difinity staff set that state. It then has 90 days to use the self-service export. At the end of that window, the organisation's data is erased unless a legal hold pauses erasure. Backups containing erased data age out seven days later.

The stand-in-to-original mapping is held apart from the trail and envelope-encrypted. A data key encrypts the values; one Platform AWS KMS key, configured for automatic rotation, protects the data key. Only a person with the PII:view permission can ask the server to reveal a mapped value, and every reveal is logged. The trail is therefore pseudonymised, not anonymous, while a revealable mapping exists.

Values masked with asterisks have no revealable mapping and cannot be revealed.

Metering and credits

Two units, never mixed

Every billable unit of work is recorded once, and the record says whose key paid for it.

Difinity funds model and tool work on credits. Where an organisation's own provider keys are enabled for it, that work is billed by the provider instead. One usage record never carries both.

Which one applies

Credits are how the work Difinity funds is metered. Work that runs on an organisation's own provider keys is priced by that provider, in the provider's own currency, and carries no credit charge.

Neither arrangement falls back to the other. A provider that cannot be paid for is left out, and if none can be paid for, the turn is refused rather than quietly charged the other way.

Tool calls are metered separately from model calls. A free tool is recorded at a zero rate rather than not recorded at all.

A request the guardrails refuse never reaches a provider and is not charged.

Running out

An organisation with no credit account is not limited by credits.

Where there is a credit account and it reaches zero, the request returns HTTP 402 with the error code OUT_OF_CREDITS. It applies to the chat and agent workspace only, and it is checked before the model call and before a stream opens. An application calling the API is not refused on credits.

Credits are a prepaid ledger, and a balance is the sum of its entries. Usage is settled shortly after each run, so runs happening at the same time can spend past a balance by the amount in flight.

Invoices are issued monthly through Stripe.

This page publishes no credit price and no credit-to-token conversion. Rates are set in the applicable order form or enterprise agreement.

Where your data lives

Hosting regions

Difinity offers hosting in AWS regions in Australia (Sydney), the European Union (Frankfurt) and the United States. The region for an organisation is agreed in the order form, and the parts Difinity operates for that organisation run there: the Platform API, Flow, the tool gateway, the workspace and the Difinity-hosted models.

The sub-processors page lists the vendors involved in delivering the service itself, and the region each one works in.

Difinity's own account and website data, including sign-in, is processed where those vendors host it. Those vendors are named in the privacy policy and the cookie policy, not on the sub-processors page.

Where it is stored

Service data sits in Aurora PostgreSQL 16.14 in private subnets with no public access, encrypted at rest with an AWS KMS key managed by Difinity, with 7-day backups and point-in-time recovery.

Attachment bytes sit in object storage under one prefix, and are removed 28 days after upload.

Provider keys and connector credentials sit in AWS Secrets Manager, never in the database. Sign-in is provided by Supabase Auth.

Credentials

Provider keys and connector credentials are write-only. They can be set, replaced and destroyed. No endpoint returns a stored credential, on the Platform API or on the tool gateway, and no screen can show one.

Removing a connector destroys every credential held for it, including the accounts individual people connected through it.

Model providers, and Difinity-hosted models

The configured model providers are OpenAI, Anthropic, Google and Grok (xAI).

Difinity-hosted models do the platform's own work. Personal-information detection, the tool-call judge and compliance drafting run on AWS Bedrock. Subjective evaluation and the requirement router run on Amazon SageMaker.

The periods that are set

The run trail defaults to seven years with a six-month minimum. Transcripts default to while subscribed or can use a set number of days. Chat attachments expire after 28 days. An agent run waiting for an approval expires after seven days. Database backups are kept for seven days with point-in-time recovery. A legal hold pauses expiry and erasure. A Terminated organisation has 90 days to use the self-service export before its data is erased; backups containing erased data age out seven days later.

No period is set for the run trail or the transcript.

Difinity staff access

Authorised Difinity staff use a separate, role-gated administration surface. They can administer organisations and people, billing, credits and plans, and view aggregate usage counts.

That surface does not give staff access to conversations, transcripts, run trails, agent definitions, prompts or uploaded files. Staff do not sign in as one of an organisation's people. There is no impersonation.

Every staff action on an organisation is written to an immutable privileged-access log that identifies the staff member and organisation. The application cannot update or delete these entries. The log is available to authorised Difinity staff and is not currently visible through customer-facing interfaces.

The data processing agreement sets out the contractual side of all of this.

Compliance evidence

Governed run records can contribute operational evidence to wider EU AI Act, ISO/IEC 42001, risk, and audit processes. Difinity does not determine that an organisation or AI system is compliant, and it does not provide ISO/IEC 42001 certification.

Legal, risk, security and compliance owners remain responsible for determining which obligations apply and how platform evidence is used within the organisation's broader programme.

Planning a first deployment

Start with one bounded job, and make the operating boundary explicit before anything is switched on. Start smaller than feels useful. Settle the commercial side with Difinity first: model and tool work runs on credits Difinity funds, unless an organisation's own provider keys are enabled for it.

Define the job

Write down the business outcome, the accountable owner, the people who will use it, the systems it touches, the data involved, and the actions it needs to take.

Decide the authority

Decide which actions the agent may complete on its own, which go to a person, and which it should not have at all. Bind only the tools the job needs, and write a per-tool policy for each tool that can change something.

Agree the evidence

Agree what security, risk, audit and business owners need to read after a run. Check that the connected systems provide the events that record needs.

Validate before expanding

Test the ordinary path, a policy refusal, a document with personal information in it, a system that is down, and an approval nobody answers. Read the resulting trail before giving the agent more to do.

Where to go next