Ethics & Responsible Use
Your Data, Their Model: How AI Companies Use What You Type
Chat history, training data and retention logs are three different things, and the defaults differ by platform and plan. Here is how to find out what applies to you.
Updated 9/13/2026
"Is it training on my conversations?" is the right question asked slightly wrong. There are at least four separate things happening to your text, and they are governed by different settings.
The four buckets
Conversation history. Stored so you can scroll back. Deleting a chat usually removes it from your view; it may persist in backups for a period.
Model improvement. Whether your content can be used to train or tune future models. This is the switch most people mean, and it is the one most often on by default for free consumer tiers.
Human review. Whether staff or contractors can read samples, typically for safety and quality. This exists on most platforms in some form, usually on a small sampled basis.
Abuse and legal retention. Logs kept for a fixed window to investigate misuse — often retained even when you have opted out of training, and sometimes extended by litigation holds.
Confusing these is how people end up genuinely surprised. Turning off training does not mean nothing is stored.
The pattern across the industry
Specifics change often enough that quoting them here would go stale, so treat this as the shape rather than the detail. Consumer free tiers are the most permissive for model improvement and usually offer an opt-out. Paid consumer tiers vary. Business, enterprise and API tiers are typically the strictest: no training on customer content by default, contractual retention limits, and admin controls. Temporary or incognito chat modes generally exclude a conversation from history and training while still permitting short-term safety retention.
Every major platform publishes this. Read the actual pages for the tools you use — OpenAI, Claude, Gemini, Grok — and check the settings screen rather than trusting a summary, including this one.
The realistic threat model
Most people are not worried about a lab reading their poem. The concrete risks are narrower and more mundane.
Pasting a client contract, patient detail, unreleased financials or someone else's personal data into a consumer chat is a disclosure to a third party, and it may breach obligations you have already signed regardless of what the platform does next. Sensitive text lodged in your own account history is exposed by any future account compromise. And detail volunteered across many conversations builds a profile richer than anything you would knowingly hand over in one go.
Model memorisation of specific inputs is possible in principle and rare in practice for ordinary text — but it is not the main hazard, and treating it as such distracts from the boring ones that actually bite.
Practical hygiene
Opt out of training where you can, and do not treat that as a substitute for judgement. Use enterprise or API access for anything client-confidential. Redact names, account numbers and identifiers by habit — models do not need them to help you. Use temporary chats for one-off sensitive questions. Clear history you no longer need. And if you build products on these APIs, tell your own users which provider processes their input, because they have the same right to know that you do.
The open argument
There is a legitimate debate here too. One side argues consented data is what makes models better for everyone, and that opt-out with clear disclosure is a reasonable bargain. Another argues meaningful consent is impossible when the disclosure is a link buried in onboarding, and that training data should be opt-in by law. Regulators in different jurisdictions are landing in different places — see AI regulation around the world.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.