The line that matters is consumer plan versus business plan
Every major vendor draws the same line. Consumer plans train on your content unless you turn it off. Business plans do not train on it at all.
OpenAI's help center states that ChatGPT Free, Go, Plus, Pro and Codex train on customer content by default, while Business, Enterprise and the API do not. For Business, OpenAI's documentation says "business data is excluded from training by default and is encrypted in transit and at rest".
Anthropic splits the same way. Its documentation says Anthropic "will train new models using data from Free, Pro, and Max accounts when this setting is on", and that it "does not train generative models using code or prompts sent under commercial terms".
Google's consumer Gemini app processes chats "to provide, develop, and improve its services (including training generative AI models)" when Keep Activity is on.
| Plan | Trained on by default? |
|---|---|
| ChatGPT Free / Go / Plus / Pro | Yes — opt out in settings |
| ChatGPT Business / Enterprise / API | No |
| Claude Free / Pro / Max | Yes when the setting is on |
| Claude Team / Enterprise / API | No |
| Gemini consumer app | Yes when Keep Activity is on |
| Google Workspace with Gemini | Governed by your Workspace agreement |
How long each vendor keeps what you type
Training and retention are different questions. A vendor can decline to train on your data and still hold it for years.
Anthropic publishes the clearest numbers. Consumer users who allow data use for model improvement get a "5-year retention period". Those who do not get "a 30-day retention period". Commercial users get 30 days as standard.
OpenAI states that "Temporary Chats are deleted from our systems after 30 days".
Google's policy is the most detailed. With Keep Activity on, chats auto-delete after 18 months by default, adjustable to 3 or 36 months. With it off, chats are kept for 72 hours. Separately, human reviewers see a subset of chats, disconnected from your account, retained "for up to three years".
Your automation tools retain data too. Zapier holds Zap history for 29 to 69 days by default.
Health data, and what a BAA actually requires
If you are a covered entity under HIPAA, the plan tier decides whether you may use the tool at all.
Anthropic is explicit: HIPAA configuration "is available for Enterprise plans only", and "Team plans and individual plans (Free, Pro, and Max) can't enable HIPAA".
OpenAI publishes a separate list of HIPAA-eligible products and a process for requesting a BAA. Consumer tiers do not qualify.
A signed BAA is a document, not a setting. Get it in writing before any protected health information goes near the tool.
Nine steps to make a US small business safe this week
- Decide which of your data is actually sensitive — client names, health records, financials, case files, employee records.
- Buy business-tier seats for anyone who touches it.
- If anyone stays on a consumer plan, turn training off. In ChatGPT: Settings → Data Controls → switch off "Improve the model for everyone".
- Tell staff to stop using thumbs-up and thumbs-down on sensitive chats. Even with training off, OpenAI's policy says that if you give feedback "the entire conversation associated with that feedback may be used to train our models".
- If you handle PHI, request the BAA before uploading anything.
- Write a one-page rule listing what may never be pasted into any AI tool.
- Ban personal accounts for work. This is the largest real gap and it is invisible on your invoice.
- Check your automation tools' retention windows, not just the model vendor's.
- Review every six months. These policies changed materially in 2025 and 2026.
Five things "we don't train on your data" does not mean
It does not mean no human reads it. Google states a subset of Gemini chats are reviewed by humans and retained up to three years.
It does not mean deleted is gone. In the New York Times litigation, a court order required OpenAI to preserve output logs that would otherwise have been deleted. The forward-looking obligation was later lifted, while data already preserved remained subject to the case. Litigation overrides a retention policy.
It does not mean one vendor. If you use Zapier or Make to send data to a model, at least two companies now hold it.
It does not cover shadow use. A US Chamber of Commerce Foundation survey by Ipsos, fielded 8–11 May 2026 with 1,070 respondents, found 50% of small business workers use AI at work while only about one in ten were offered formal training. Your staff are already using these tools. The question is on whose account.
It does not mean your output is protected. The US Copyright Office concluded in Part 2 of its AI report that prompting alone does not establish human authorship.
What most articles get wrong about this
They tell you ChatGPT is unsafe, or that it is safe. Both are wrong, because the tool is not the unit of analysis. The same company, on the same day, is training on one employee's Plus account and contractually barred from training on the next employee's Business seat. The fix is a $20 seat and a settings toggle, not a ban.
The second error is assuming the toggle protects you retroactively. Turning training off in Claude moves you to a 30-day window going forward. It does not undo the five-year window that applied while it was on.
This page explains what the published policies say. It is not legal advice.
Where these figures come from
- OpenAI: how your data is used · ChatGPT Business privacy · Data Controls FAQ
- OpenAI HIPAA-eligible products · Requesting a BAA
- Anthropic data usage policy · HIPAA-ready Enterprise plans
- Google: how Gemini Apps uses your data · Workspace HIPAA compliance
- Zapier data retention
- US Chamber of Commerce Foundation, Main Street AI Monitor
- US Copyright Office, Copyright and AI · FTC Operation AI Comply
People also ask
- How much does it cost to use AI in a small business?
- Do I have to tell customers when a reply is written by AI?
- What can AI actually do for a small business, and what can it not do?








