Which AI Tools Train on Your Data by Default? (Comparison)
Comparison dated 2026-08-08: which AI tools (ChatGPT, Claude, Gemini…) train on your chats by default, and where to find the opt-out.
Which AI tools train on your data by default? The honest answer: it depends on the tool and the plan. On consumer tiers — free, Plus, Pro — several assistants use your conversations to improve their models by default, with an opt-out you switch on by hand. On API and Enterprise offerings, training is usually off by default. The table below is a snapshot dated 2026-08-08, with each row attributed to its vendor. These policies change often: always check the official page before you conclude.
The short answer, before the table
First, hold on to the dividing line. It explains 90% of cases.
- Consumer (free, Plus, Pro): training by default is common, with an opt-out to switch on manually.
- API and Enterprise: training is usually off by default, framed by a data processing addendum (DPA).
- The opt-out often only applies going forward: it doesn’t erase data already used.
- An employee’s personal account follows consumer rules, not the guarantees of an enterprise plan.
Attributed comparison, as of 2026-08-08
Each cell reflects the policy stated by the vendor, as read on 2026-08-08. When a detail can’t be verified, it is marked « to verify ». That’s not a judgment: it’s a snapshot that can move.
| Tool | Consumer: training by default? | Opt-out? | API / Enterprise |
|---|---|---|---|
| ChatGPT (OpenAI) | Yes (Free/Plus/Pro), per the OpenAI help center | Yes (« Improve the model for everyone » setting) | Not by default (API, enterprise offerings) |
| Claude (Anthropic) | Yes since 28/08/2025 (Free/Pro/Max), per Anthropic | Yes (Privacy Settings; exception: chats flagged for safety review) | Not by default (API, Team/Enterprise) |
| Gemini (Google) | Yes, with human review via « Gemini Apps Activity », per Google | Yes (turn off Activity; also trims history) | Stronger protections on Workspace/enterprise |
| Microsoft 365 Copilot | Enterprise: No (Enterprise Data Protection), per Microsoft | n/a (foundation models not trained) | No (prompts/responses not used for training) |
| Le Chat (Mistral AI) | Yes (free account), per Mistral AI | Yes (privacy setting) | Not by default (Team, Enterprise, API) |
| Perplexity | Yes (Free/Pro/Max), per Perplexity | Yes (account preferences) | No (Enterprise is never used) |
| DeepSeek | Yes, with storage in China (Chinese law), per DeepSeek | Yes | To verify, depending on plan |
| Grok (xAI) | Yes (on by default), per xAI | Yes (two dedicated settings) | Separate enterprise terms |
| GitHub Copilot (IDE) | Yes since 24/04/2026 (Free/Pro/Pro+), per GitHub | Yes | No (Enterprise/organization accounts, governed by the DPA) |
Why it changes, and how to check
Vendors adjust their consumer terms regularly. The CNIL offers a useful frame for reading these policies. Per the CNIL, the GDPR applies in many cases to models trained on personal data, because of their capacity to memorize. It adds that people must be informed when their data serves training. Here’s how to check for yourself, in four steps.
- 1Open the tool’s privacy page and find the « training » or « model improvement » section.
- 2Identify your exact plan: a consumer setting often differs from an API or enterprise contract.
- 3Locate the opt-out setting and note whether it only applies going forward.
- 4Date your reading and schedule a recheck, since the policy can change without a visible notice.
The control that depends on no vendor
There is one control that doesn’t change when the terms change: anonymize the identifiers in your prompt before you send it. If the AI only receives a token instead of a name or an IBAN, the question « does it train on this? » loses most of its stakes. It’s a lever you keep in hand, whatever the vendor’s policy. It reduces your data’s exposure; it doesn’t remove it, and it doesn’t replace compliance work.
ONYRI Sanitize applies exactly this principle. The extension grafts onto AI chat sites — ChatGPT, Claude, Gemini, Le Chat, Perplexity, Copilot web, DeepSeek, Grok. It acts in your browser, before you send. It doesn’t cover an IDE’s autocomplete, nor a tool’s internal AI calls: it’s a layer on the chat prompt, nothing more.
In short, the training default varies by tool and by plan, and it evolves. On consumer tiers it’s often on; on API and enterprise, rarely. Check each vendor’s page, because this comparison is dated. The defense that depends on no policy stays the same: anonymize the identifiers before sending. ONYRI detects, replaces with reversible tokens, then restores the answer in your browser. Reversible tokens remain pseudonymization: this reduces exposure, without making your data « anonymous under the GDPR ».
Frequently asked questions
- Which AI tools train on my data by default (ChatGPT, Claude, Gemini…)?
- It depends on the tool and the plan, as of 2026-08-08. On consumer tiers, several assistants train by default with an opt-out. Examples: per Anthropic for Claude (Free/Pro/Max), per Mistral AI for the free Le Chat, per Perplexity for its Free/Pro/Max offerings. On API and enterprise, training is usually off by default. Always check the vendor’s page in force.
- Does the opt-out erase data already used for training?
- Usually not. The opt-out often only applies to future exchanges and doesn’t erase what has already been used. Some settings have exceptions: per Anthropic, chats flagged for safety review may still be used despite the opt-out. That’s why anonymizing before sending beats relying on a retroactive setting.
- Does anonymizing with ONYRI make me GDPR-compliant?
- No, not on its own. Reversible tokens are pseudonymization: your data stays personal under the GDPR. The benefit is concrete: your real values don’t leave your browser, which strongly reduces exposure and serves the minimization principle, whatever the vendor’s training policy.
Sources & references
- Updates to Consumer Terms and Privacy Policy (consumer Free/Pro/Max moved to training by default, with opt-out, 28 August 2025) — Anthropic
- Data, Privacy, and Security for Microsoft 365 Copilot (prompts, responses and Microsoft Graph data not used to train foundation models) — Microsoft
- AI and GDPR: new recommendations (the GDPR applies to models trained on personal data; informing individuals) — CNIL
Keep your sensitive data in your browser
ONYRI Sanitize detects and masks your sensitive data before it reaches the AI, then restores the answer — from names to API keys.