Does AI Train on Your Data? 12 Tool Policies Checked

Does AI train on your data? For most tools marketers and creators use daily, yes, by default. We read the official privacy and data control pages for 12 AI tools on 19 September 2026. In 8 of the 12, content from a normal consumer account is used to improve or train models unless you change a setting. In 2 the stated default is no training. One was clear only in part, and one did not state a default.

Almost every provider gives you a switch. 11 of the 12 name a setting you can turn off, and 11 say business, enterprise or API data is excluded by default. The gap is rarely the control. It is that nobody on the team knows the control exists.

This records what those pages said on the day we read them. It is not legal advice and not an accusation against any company. Policies change, so open the linked page before deciding anything.

Key takeaways

  • 8 of 12 tools use consumer account content to improve or train models by default.
  • 2 stated the opposite default: no training unless you switch it on.
  • 11 of 12 name a setting you can change, and 11 exclude business or API data by default.
  • Only 6 of 12 said anything specific about retention after deletion or opt out.
  • Opt outs are forward looking: they cover what you submit next, not what you already sent.
  • Checked 19 September 2026 on official pages only. Verify before you rely on any row.

Does AI train on your data by default?

For most consumer accounts, yes. Across 12 tools checked on 19 September 2026, 8 used free consumer content to improve or train models by default, 2 did not, 1 did so for product improvement but not generative training, and 1 did not state a default. Business and API tiers were excluded in 11 of 12.

The pattern is consistent enough to be a working rule. The account you created in two minutes with a personal email is the one most likely to feed a training pipeline. The account your company pays for under a contract is the one carved out. That is not a scandal, since providers disclose it, but it is easy to miss when you are working fast.

Two tools ran the other way. Anthropic states that consumer Claude chats are not used for model training unless you turn model improvement on, with safety flagged conversations as an exception, per its model training help page. Notion states that “By default, Notion and its AI Subprocessors do not use Customer Data to train any models” on its AI security practices page.

Why does this matter for client work and creator IP?

Two risks sit behind the setting. One is contractual: client agreements often restrict sharing confidential material with third parties, and a default on training setting can make a tool an undeclared third party. The other is ownership: creators paste scripts, voice samples and unpublished footage into tools whose terms allow model improvement.

Start with the contract risk, because it is the one that ends relationships. A media plan or an unreleased product name is the client’s confidential information. If your agreement limits disclosure to named subprocessors and someone drops the file into a personal AI account, the disclosure question is real even if no harm follows.

The creator side is quieter and just as concrete. A voice sample carries your brand; a rough cut is unpublished work. The ElevenLabs data use article says data you provide improves audio models for everyone, with an opt out open to anyone, and Descript says you can opt out of projects improving the service in its account data and privacy article. Both are easy to miss at midnight. What you type is an asset, the logic behind our guide to prompt engineering for marketers.

What exactly did we check, and how?

We read 21 official pages across 12 AI tools on 19 September 2026: privacy policies, data control help articles and trust centre pages published by the providers. For each tool we recorded the free consumer default, the paid default, whether business or API data is excluded, the setting name and location, and any stated retention.

Four rules kept the result honest:

  • Official sources only. No news, forums or third party summaries.
  • One short quoted phrase per tool at most. The rest is our own description.
  • “Not clearly stated” is a valid answer. Where a page gave no default or retention period, we recorded that instead of guessing.
  • Consumer view, one date. We read the pages a signed up user reaches, and every row carries the same check date.

Raw rows and page URLs sit in one file, so every count traces to a source.

Which AI tools train on your content by default?

Here is the full comparison: the free consumer default, the paid or business position, how to opt out, and the date checked. Settings and names change, so open the linked page before acting. Every entry reflects what the provider published on 19 September 2026.

Tool Free consumer default Paid or business position How to opt out Checked
ChatGPT (OpenAI) Yes, used to improve models Plus and Pro same. Team, Enterprise, Edu and API excluded Settings, Data Controls, Improve the model for everyone 19 Sep 2026
Google Gemini Apps Yes, when Keep Activity is on Consumer paid same. Workspace data excluded without permission Gemini Apps Activity, turn off Keep Activity 19 Sep 2026
Claude (Anthropic) No, unless you switch it on Pro and Max same. Commercial and API under separate terms Settings, Privacy, leave model improvement off 19 Sep 2026
Perplexity Yes, on by default Pro and Max same. Enterprise never used for training Account settings, Preferences, AI data retention 19 Sep 2026
Microsoft Copilot (consumer) Not clearly stated Does not cover work or school accounts Settings, Privacy, Training on conversation activity 19 Sep 2026
Canva Yes, a listed use of content Teams, Business, Enterprise and Education excluded Settings, Privacy, privacy preferences page 19 Sep 2026
CapCut Yes, a listed purpose Not clearly stated for paid or business tiers No named training opt out found 19 Sep 2026
Grammarly Yes, on by default for Free and Premium Sales bought, Enterprise and Education off by default Account settings, Privacy, Product Improvement and Training 19 Sep 2026
Notion AI No, stated as not used to train models Same default across plans Settings, Notion AI, Share data to improve Notion AI 19 Sep 2026
ElevenLabs Yes, used to improve audio models Enterprise data not trained on by default Profile, Terms and privacy, Data use 19 Sep 2026
Descript Yes, projects may improve the service Enterprise drives have no toggle, sharing disabled App settings, Profile, Share data with Descript 19 Sep 2026
Adobe Creative Cloud Partly: analysis on, generative training stated as not done Business, team and school profiles opted out account.adobe.com, Privacy, Content analysis 19 Sep 2026

Two rows need a note. Adobe splits the question: its content analysis FAQ describes analysis for product improvement as on for personal accounts while stating content is not analysed to train generative AI models unless you submit to Adobe Stock. CapCut was the only tool with no named training opt out on the pages we read, although its policy lists training machine learning models as a purpose.

What patterns showed up across the 12 policies?

Five patterns repeated. Consumer defaults lean towards training while business tiers are carved out. Paying as an individual rarely changes the answer. Opt outs are forward looking only. Retention after deletion is the least documented area. And every provider uses a different word for the same switch.

1. The business tier is the real privacy tier. 11 of 12 tools exclude business, enterprise or API data by default. If a workflow touches client material, the account type matters more than the tool.

2. Paying as an individual changes little. The free and paid consumer default matched in 10 of the 12 rows. A personal upgrade buys capacity and features, rarely a different data position.

3. Opt outs start when you flip them. Perplexity states opt outs apply only to data collected after the opt out date, per its data collection article. ElevenLabs states the same limit. That makes the setting a day one task, not an after the fact fix.

4. Retention is the weakest disclosure. Only 6 of 12 said anything specific about content after deletion or opt out. Anthropic describes deleted chats leaving back end storage within 30 days, and training pipeline data kept up to 5 years in de-identified form when model improvement is on, per its data storage article. Google states chats picked for human review are kept up to three years and survive deleting your activity, per the Gemini Apps Activity page. Microsoft states an 18 month history window.

5. Every provider uses a different word. Improve the model for everyone. Keep Activity. Model improvement. AI data retention. Training on conversation activity. Product Improvement and Training. Content analysis. No shared vocabulary, which is why teams miss it. If you are comparing output quality too, see our side by side on ChatGPT vs Gemini vs Perplexity is a useful companion.

What should your team AI data policy say?

Six steps cover most of the risk: inventory the tools, move client work onto business accounts, turn training off on every personal account, set a red list of content that never goes into a consumer tool, record the check date, and re-check quarterly.

Step What to do Who owns it How often
1. Inventory List every AI tool anyone uses, including ones nobody approved. Ops or team lead Once, then quarterly
2. Move client work Put workflows touching client material on a business plan, where 11 of 12 providers exclude data by default. Ops plus finance Once, then at renewal
3. Flip the switches Open the data control page for every personal account, turn training off, screenshot it. Each individual Day one of any account
4. Write a red list Name what never goes into a consumer tool: unreleased pricing, personal data, legal drafts, raw client footage. Team lead with legal input Once, review yearly
5. Record the date Log the tool, the setting changed and the date. This is what you show a client who asks. Ops Each check
6. Re-check quarterly Set a reminder. Several pages here had been revised within the last year. Ops Every quarter

Formula: Audit minutes = personal accounts to check x 3 minutes

Example: say your team has 12 people. Six use an AI assistant daily, four a design tool with AI features, two an audio or video editor. That is 12 accounts in scope. Five sit on a company managed business plan, excluded by default in 11 of 12 cases. The other 7 are personal, so 7 x 3 = 21 minutes of clicking. Add 30 minutes to write and circulate the policy and the job fits in an hour.

Keep it to one page. A nine page policy gets read once and ignored. One naming five things that never get pasted into a free tab gets followed. Our list of free AI tools for creators and the collection of AI prompts for creators are practical starting points.

What should you do about client confidentiality?

Three moves: check what your contract says about third parties and subprocessors, use a business tier account for anything it covers, and tell the client which tools you use before they ask. Do not rely on an opt out setting alone to satisfy a confidentiality clause.

Many agency contracts name approved subprocessors or require notice before confidential material goes to a new vendor. An AI tool is a vendor. Switching off model improvement does not change that content was transmitted, stored and processed. The setting reduces one risk; it does not answer the contract question.

A short disclosure beats silence: these are the tools we use, they sit on business plans, training is excluded by default or by a setting we turned off, and here is the date we verified it. Where material is genuinely sensitive, the safest answer is the boring one: summarise or anonymise it. Teams that already think carefully about first party data and clean rooms find this an easy habit, and our digital media planning glossary defines the surrounding terms in plain language.

What this means for marketers and creators

What this means for marketers and creators

For marketers: the account type is the control that matters. Move any workflow touching client briefs, pricing, customer lists or unreleased creative onto a business plan, because 11 of 12 providers exclude that tier by default. Then turn training off on every personal account still used for work, and add the tool list and check date to client onboarding. Describe it as the provider default on a stated date, not a guarantee.

For creators: your voice, your face, your rough cuts and your scripts are the assets. Check audio and video tools first, because those take the richest material. ElevenLabs and Descript both describe an opt out any user can set, applying to what you submit after you flip it, so do it before the next upload. If you take brand deals, expect sponsors to ask which tools touched the footage. Having a one line answer ready is a small professional advantage.

What we did not check

This is a reading of published pages, not an audit of systems. We did not test what the tools do, read enterprise contracts, compare regional versions, or cover tools outside this set of 12. Treat it as a map of what providers say, not proof of what happens.

  • No technical verification. We did not inspect traffic or test outputs, and nothing here claims any provider behaves differently from its policy.
  • No contracts or addenda. Negotiated agreements often go further than the public page.
  • Region matters. Defaults can differ by jurisdiction. We read the general English pages.
  • Twelve tools is a small set. This covers common ones, not the market.
  • One date, no legal advice. Everything reflects 19 September 2026, and obligations differ by country and client.

The method is repeatable. Search the provider site for a privacy or data controls page, look for the words train or model improvement, note the default, the setting name and location and the date, and write “not clearly stated” when the page does not say. That habit matches the spirit of Google’s own guidance on AI generated content: describe what is known, not what is assumed.

Frequently asked questions

Does ChatGPT train on your data by default?

On free and paid consumer accounts OpenAI states content is used to help improve models unless you turn the setting off in Settings under Data Controls. Business tiers including Team, Enterprise, Edu and the API are stated as excluded by default. Checked 19 September 2026.

Which AI tools do not train on your content by default?

In our set of 12, two stated a no training default: Claude, where consumer chats are not used for training unless you switch model improvement on, and Notion, which states that it and its AI subprocessors do not use customer data to train models.

Does paying for a subscription stop AI training?

Usually not on an individual plan. The free and paid consumer default matched in 10 of the 12 tools checked. The tier that changed the answer was business, team, enterprise or API, which 11 of 12 providers described as excluded by default.

Where is the setting to turn off AI training?

It differs by tool and there is no common name. Look under Settings for data controls, privacy, activity, model improvement or product improvement. Examples here include Data Controls in ChatGPT, Gemini Apps Activity, Preferences in Perplexity and Privacy in Copilot.

Does opting out delete data the tool already used?

Generally no. Perplexity states opt outs apply only to data collected after the opt out date, and ElevenLabs describes the same forward looking limit. Treat the setting as something to switch on before a project starts, not a way to claw back earlier content.

Is it safe to put client confidential work into an AI tool?

That depends on your contract, not the tool alone. Check whether your agreement restricts sharing with third parties or requires named subprocessors, use a business tier account for covered work, and tell the client which tools you use. For sensitive material, summarise or anonymise.

Next steps: block 30 minutes this week, run the six step policy on your stack and write the date next to each tool. For more on putting AI to work safely, read what an AI agent actually is and browse the free TechMachaw tools. Join the free TechMachaw newsletter to get the next study as soon as it publishes.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top