Home/Learn/How to opt out of AI training on your company data
Tutorial

How to opt out of AI training on your company data

How to exclude company data from AI training: check the tier and contract, set retention terms and add a technical control.

What opting out does, and what it does not do

Training exclusion is a promise about one use of your data. It is not a retention limit, it is not a bar on abuse monitoring, and it does nothing about an employee pasting a customer list into a personal account at home. OWASP catalogues that gap as sensitive information disclosure in its Top 10 for LLM Applications, and the reason it stays on the list is that the contractual control and the technical control solve different halves of the problem. Treat the opt-out as necessary and clearly insufficient. The question that follows it is what leaves your boundary in the first place.

Step 1. Find out which tier each team is actually on

The split that matters runs between business access and consumer access, not between vendors. Business and API access is where the training exclusions live, and both OpenAI and Anthropic state the exclusion in their own words on their own pages. A personal subscription sits under a different document with different defaults, which is why the shadow-AI question and the training question turn out to be the same question. Pull the expense reports and the identity provider logs before you accept anyone's answer about which tier their team uses. Here is what each provider published as of 2 September 2026, and every line needs confirming against your own contract rather than this page.

  • OpenAI: data sent to the API has not been used to train or improve OpenAI models since 1 March 2023, unless you explicitly opt in
  • OpenAI retention: abuse-monitoring logs are generated for API usage and kept for up to 30 days by default. That is a log-retention figure, not a blanket statement about your content
  • OpenAI options: approved customers can enable Zero Data Retention, which keeps customer content out of those abuse-monitoring logs entirely, or Modified Abuse Monitoring, which does the same while keeping full API functionality
  • Anthropic training: inputs and outputs from commercial products, including the API and Claude for Work, are excluded from model training by default, in Anthropic's own words
  • Anthropic retention: API inputs and outputs are automatically deleted from Anthropic's backend within 30 days of receipt or generation
  • Google: Gemini does not use your prompts or its responses as data to train its models, in Google's own words. Its data governance page publishes no retention period in days, so that one is a question for your contract rather than a figure you can look up

Step 2. Get it in the agreement, not in a help-centre article

A published policy is a statement of current practice and can change with a product update. The clause in your agreement is the thing you can hold someone to. Ask for each of the points below in writing, and ask specifically what survives a change to the provider's standard terms, because that answer tells you how much the rest of it is worth.

  • The training-exclusion clause itself, naming inputs and outputs, plus anything derived from them
  • The retention period, and whether a zero-retention arrangement is available for your volume
  • What is carved out for abuse monitoring, who can read it, and for how long it is kept
  • The sub-processor list, and the notice you get before it changes
  • Which region processes the data, and what happens to it at the end of the contract

Step 3. Turn off the settings that survive the contract

Contract terms do not reach account settings that individual people control. Check the administrator console for any workspace-level improvement or model-training toggle and record its state with a screenshot and a date. Check whether people can sign in with a personal email address alongside the corporate one, because the same interface with a different account can carry different terms. Then look at the surface nobody owns: browser extensions, meeting note takers, code assistants and the AI features a vendor quietly enabled inside a product you already licensed.

Step 4. Add a control on your side, because a promise is not one

This is where the older version of this page told you to route traffic through a gateway. That is not how Difinity.ai works, so here is the accurate version. Difinity has four configured model providers: OpenAI, Anthropic, Google and Grok (xAI). Difinity-hosted models do the platform's own work, with personal-information detection, the tool-call judge and compliance drafting running on AWS Bedrock, and subjective evaluation and the requirement router on Amazon SageMaker, all in the region the organisation agreed in its order form. The control that matters happens before any of that. Where a use case is configured to detect personal information, the message is checked and detected values are replaced, and the real values are restored only in the last step before an action leaves through the tool gateway. The model call sees the redacted form for the values detection found, which is a technical answer to the question a training clause can only answer contractually. Provider keys and connector credentials are write-only: no endpoint returns a stored credential through customer-facing interfaces, and no screen can show one.

Step 5. Verify it, then keep verifying it

An opt-out you cannot evidence is a belief. Pick a real conversation from last week and check which model answered it, because the answer carries the model that produced it. Read the run trail for the same turn: it records the kind of value that was replaced, never the value itself, which is the property that lets you show a reviewer what was protected without handing them the thing you protected. Put a calendar trigger on each provider's terms. These figures move, and third-party summaries of them go stale faster than the terms themselves, which is why every number above is dated and taken from the provider's own page rather than from someone's write-up of it.

Where this gets you, and where it stops

An opt-out is forward-looking. It says nothing about data already processed under earlier terms, and no provider offers to unlearn a model. Abuse monitoring is usually carved out of every commitment on this page, so assume a human at the provider can read flagged content unless your contract says otherwise. And none of this covers the traffic that never touched the governed path: a personal account, an unmanaged browser extension, a vendor feature switched on last quarter. Every provider line above comes from that vendor's own current page, checked on 2 September 2026. Two of the three publish a retention period and Google does not, which is worth noticing rather than assuming a number exists somewhere. None of them are the legal text: a product page tells you what a vendor currently does, your agreement tells you what it owes you, and only the second one is enforceable.

Frequently asked questions

Does the OpenAI API use my data for training?

OpenAI states that data sent to the API has not been used to train or improve its models since 1 March 2023, unless you explicitly opt in. Abuse-monitoring logs are kept for up to 30 days by default, which is a separate matter from training, and approved customers can keep customer content out of those logs. Consumer ChatGPT accounts are governed by different terms, which is where most organisations find their real exposure.

Does Anthropic train on Claude API data?

No. Anthropic states that inputs and outputs from its commercial products, the API included, are excluded from model training by default, and that API inputs and outputs are automatically deleted from its backend within 30 days. A personal Claude.ai subscription sits under separate consumer terms, so read those before assuming the commercial position covers it.

Is opting out of training enough to protect confidential data?

No. Training exclusion is a promise about one use of the data. It does not limit retention, does not cover abuse monitoring, and does not stop anyone pasting confidential text into an account your organisation does not control.

What should be in the contract rather than a policy page?

The training-exclusion clause, the retention period, any zero-retention arrangement, the abuse-monitoring carve-out, the sub-processor list with a notice period, and the processing region. A published policy can change with a product update. A clause cannot.

How does Difinity keep company data out of a provider's model?

Where a use case is configured to detect personal information, values are replaced before the model call and restored only in the last step before an action leaves through the tool gateway, so the model works on redacted text. That is a technical control alongside your contractual one, not a replacement for it.

Sources and further reading

Have an agent that needs production authority?

Read the PII redaction guide