SELINA.ai
Sign in

Is My Data Safe with AI? What Actually Happens to Your Inputs, Logs, and Training Data

You typed something into a chatbot last week. Maybe it was a business idea, a medical question, a rough draft of a contract. You probably closed the tab and moved on. But the question worth pausing on is: where did that text go? Is my data safe with AI, or is it sitting on a server somewhere, feeding the next version of a model you never agreed to improve? The answer depends on architecture, not promises. This piece walks through what happens to your data across common AI tools, what courts and regulators are doing about it, and what you can actually verify before trusting a product with anything that matters.

Key Takeaways

What Happens to Your Data When You Use an AI Chatbot?

Your input leaves your device, crosses a network, and lands on a server. That much is obvious. What's less obvious is what happens next. At most major AI providers, the default behavior is to retain your conversation and use it as training material for future model versions. A 2026 review of the ten most popular AI chatbots found that nine of them train on user conversations by default. To stop this, you typically need to locate a toggle buried several layers deep in account settings, sometimes under a label that has nothing to do with training (like "improve our models" or "data sharing preferences").

Even when you find the opt-out, you're trusting the provider's implementation. There's no external audit confirming that flipping a switch actually purges your data from training pipelines already in progress. You're reading a policy page and deciding to believe it.

That same review found six of those ten chatbots had already faced fines, bans, or formal regulatory inquiries. One stores all data in a jurisdiction where the government can compel companies to hand it over on request. These aren't edge cases. They're the products most people use.

Does "Deleted" Actually Mean Deleted?

No. "Deleted" is a UI promise. It means the text disappears from your screen. It does not necessarily mean the data is gone from the provider's servers, from their backups, or from legal reach.

Here is a concrete example. In May 2025, a federal judge issued a preservation order requiring a major AI provider to retain all user chat logs, including content users believed they had already deleted. This was part of unrelated copyright litigation. The provider's own deletion mechanisms were overridden by a court order. Then in January 2026, a U.S. district court upheld orders requiring that provider to produce a sample of 20 million de-identified user chat logs to copyright plaintiffs.

Twenty million. De-identified, yes. But produced. Handed over. The data existed because it had never actually been destroyed, and legal process overrode whatever the deletion button implied.

This isn't a flaw in one company's engineering. It's structural. Once data leaves your device and sits on someone else's infrastructure, legal process (litigation holds, subpoenas, regulatory inquiries) can override the company's own deletion policy regardless of its intentions. The company might sincerely want to delete your data. A judge can say no.

This reframes the whole question. Data privacy in AI is not about what a company's policy page says. It's about where the data physically lives and who can compel access to it.

Can My AI Chat History Be Subpoenaed?

Yes. Courts have begun treating AI conversations as ordinary business records. They are discoverable. They are not privileged.

A federal court ruling found that prompts, outputs, and activity logs from AI systems can be treated as standard electronic records, subject to the same discovery rules as emails or text messages. The court also found that attorney-client privilege does not apply to conversations with an AI system, because the system is not a licensed professional. You don't have a confidential relationship with a chatbot, legally speaking.

This matters beyond big tech lawsuits. Reporting from 2026 notes that AI chat logs are already surfacing in divorce cases, employment disputes, and contract fights. If you typed it into a chatbot, opposing counsel in an unrelated matter could potentially request it. The "I have nothing to hide" instinct doesn't hold up once your offhand queries become a permanent, searchable record that a court or employer can later demand.

Does My AI Tool Use My Data for Training?

Probably, unless you've explicitly turned it off. And maybe even then.

The default across most consumer AI products is opt-out, not opt-in. A 2026 guide on stopping AI training on your data flags that following major policy shifts in late 2025, many large providers moved to models where users must actively find and disable training. The setting is often not surfaced during onboarding. You have to know it exists, know what it's called, and trust that toggling it works retroactively on data already ingested.

Enterprise customers sometimes get contractual protections. API-tier access at some providers comes with terms stating that inputs won't be used for training. But consumer-tier and free-tier products rarely offer these guarantees. If you're using the version that doesn't cost you anything, you're typically the training data.

A 2025 TELUS survey found that 57% of enterprise employees admitted to entering sensitive or high-risk information into public AI assistants. Not into enterprise-contracted, walled-off deployments. Into the same consumer chatbots that train on inputs by default. The gap between what employees do and what their employers assume they're doing is significant.

What About My Company's Data in Enterprise AI Tools?

Enterprise vendors increasingly repurpose customer content for AI training, sometimes with a policy change buried in a blog post months before the effective date. One example: a major cloud productivity suite announced in mid-2026 that it would begin using data from its products to train AI features starting in August, affecting roughly 300,000 customers. The mechanism was a policy update. If you missed the announcement, your organization's data was enrolled by default.

This is a recurring pattern. Enterprise software vendors add AI features, then adjust their data-use terms to feed those features. The AI capability is the marketing story. The expanded data-use clause is the footnote.

Are My Vendor's Privacy Promises the Whole Story?

Usually not. And this is the trap most small businesses fall into.

If you use a chatbot product built on top of a large underlying AI model, there are at least two data policies in play: your vendor's, and the model provider's. These are often different. A 2026 vendor guide highlights that many chatbot tools built on a foundation model retain and use conversations on their own terms, regardless of what the upstream model provider does. Your vendor might promise not to train on your data, while their model provider ingests everything your vendor sends through the API.

If you're evaluating an AI tool for your business, three questions cut through the noise:

  1. Who else touches this data? Not just your vendor, but whatever model provider sits underneath.
  2. Is training on by default? At every layer, not just the one with the nice UI.
  3. What's the actual retention window? Not "we delete it" but how long, under what conditions, and what overrides that (like a legal hold).

Most vendors won't volunteer clear answers to these. The ones who do are telling you something about how they think about the problem.

How Are Courts and Regulators Responding to AI Data Risks?

Slowly, unevenly, and with growing urgency.

On the regulatory side, the FTC opened a formal inquiry in September 2025 into AI companion chatbots, sending information demands to seven companies about how they measure and monitor harms to children and teens. That scrutiny has continued into 2026, but Congress has not passed comprehensive AI privacy legislation. The regulatory posture is enforcement actions and inquiries, not a coherent framework.

Public confidence tracks accordingly. Pew Research Center found in February 2026 that roughly seven in ten U.S. adults predict AI will make their personal information less secure. Just 3% think it will make information more secure. And 67% of Americans say they have little to no confidence in the government's ability to regulate AI effectively, up from 62% in 2024.

The court system is moving faster than legislators. The rulings on chat log discoverability, the preservation orders, the production of 20 million logs to copyright plaintiffs: these are creating precedent right now. The practical effect is that AI providers are being treated like any other company that holds user records. If a court wants the data, and the data exists, it gets produced.

Can GDPR Protect Me If I'm in Europe?

In theory. In practice, cross-border conflicts are emerging. A U.S. preservation order can directly conflict with the EU's "right to be forgotten." A provider legally compelled to retain data in the United States may be unable to honor a European user's deletion request. The legal frameworks haven't been reconciled, and until they are, your GDPR rights may not survive contact with U.S. litigation.

What Should I Actually Look For in an AI Product's Privacy Practices?

Ignore the adjectives. Read the architecture.

A privacy policy that says "we take your privacy seriously" tells you nothing. What matters is how the system is built. Specifically:

Where does inference happen? If your input is sent to a remote model provider, that provider's terms apply. If the provider's API terms say they don't train on API inputs, that's better than consumer-tier terms, but it's still a policy commitment, not a technical guarantee.

What's retained after inference? Some products store full conversation logs indefinitely. Others retain operational metadata for a short retention window and discard content. The difference matters enormously if a court order arrives.

Is the data encrypted, and who holds the keys? Encryption in transit (between your device and the server) is table stakes. Encryption at rest (on the server's disk) is better. But if the provider holds the decryption keys, they can still access and produce the data under legal compulsion. Client-side encryption, where you hold the key and the provider literally cannot read the content, is a fundamentally different architecture.

Can you actually verify any of this? Open-source components, published audits, and reproducible claims are worth more than policy language. A company that says "we can't read your data" should be able to explain, technically, why that's true.

How Does Selina Handle Data Differently?

We built Selina as a privacy-focused AI assistant that remembers you across conversations. Memory is adaptive and encrypted at rest. The product runs on a stack of frontier models, routed per task, accessed via API.

Different paths through the product carry different protections, and we state them separately because they're separate facts. A file uploaded into SelinaSEND is zero-knowledge, end-to-end encrypted. A standalone note written outside the AI is also zero-knowledge. A chat message is encrypted in transit and at rest, but the content was processed by a frontier model during creation, so it was never zero-knowledge. A file sent from SelinaVault is encrypted at rest, under a key you control, and cryptographically erasable. Memory is not end-to-end encrypted: a slice of each request reaches a frontier provider at inference. We keep non-content operational metadata for a short retention window.

Delete means gone. Actually gone. For the paths where we hold no key, we literally cannot read the content. For the paths where inference is involved, we minimize what's retained and make the honest limit visible rather than burying it in a footnote.

Why Does Architecture Matter More Than Policy?

Because policy is a promise. Architecture is a constraint.

A policy says "we won't look at your data." An architecture says "we can't look at your data, because we don't have the key." A policy can be changed with a blog post and a 30-day notice period. An architecture requires rebuilding the system.

The court orders we discussed earlier illustrate this perfectly. The providers who were compelled to produce chat logs weren't violating their own policies. They were complying with a legal obligation that their architecture couldn't prevent, because the data was accessible to them. If a provider physically cannot decrypt your content, a court order to produce it is moot. You can't produce what you can't read.

This is why the question "is my data safe with AI" can't be answered by reading a privacy policy. It can only be answered by understanding the system's design. Who has the keys? What's retained? What can be compelled?

What Can You Do Right Now?

A few concrete steps, none of which require technical expertise:

Check your training opt-out settings. For every AI chatbot you use, open the settings and look for anything related to model training, data sharing, or improvement programs. Turn it off. Guides exist for most major products. This takes five minutes per tool and is the single highest-leverage thing you can do today.

Treat every chatbot input as potentially permanent and discoverable. Before typing something into an AI tool, apply the same judgment you'd apply to sending an email from your work account. Could this surface in a legal dispute? Would you be comfortable with opposing counsel reading it? If not, don't type it.

Ask your vendors the three questions. Who else touches this data? Is training on by default? What's the actual retention window? If your vendor can't answer clearly, that's your answer.

Separate sensitive workflows from general-purpose AI tools. If you need AI assistance with confidential material, use a product that's architecturally built for confidentiality, not one that bolted a privacy toggle onto a system designed for data collection.

Your data is not automatically safe with AI. It can be, if the system is built for it. The difference is whether privacy is a feature toggle or a structural property of the system. Read the architecture, not the marketing.

Start a free 7-day trial, no card required.

Frequently Asked Questions

Do AI chatbots use my conversations to train their models by default?

Yes, a 2026 review found that nine out of ten popular AI chatbots train on user conversations by default. You typically have to find and manually disable a setting, often buried under an unrelated label, to stop it.

If I delete my chat history, is it really gone?

Not necessarily. 'Deleted' means the text disappears from your screen, but the data may still exist on the provider's servers or backups, and legal orders can override deletion policies entirely, as happened when a federal judge ordered a major AI provider to preserve and later produce millions of user chat logs.

Can my AI chat logs be used against me in a legal dispute?

Yes, courts have ruled that AI prompts, outputs, and activity logs are treated as ordinary discoverable business records, similar to emails, and are not protected by attorney-client privilege. Such logs have already surfaced in divorce cases, employment disputes, and contract fights.

Are enterprise AI tools safer for company data than consumer chatbots?

Not automatically. Enterprise vendors have been known to change data-use terms via policy updates, such as one cloud productivity suite that began using customer data to train AI features affecting roughly 300,000 customers, often with default enrollment unless the change was noticed.

If my AI vendor promises not to train on my data, can I trust that?

That promise may only cover part of the picture, since many chatbot tools are built on an underlying foundation model with its own separate data policy. Your vendor might not train on your data while the model provider underneath still does, so it's worth asking who else touches the data, whether training is on by default at every layer, and what the actual retention window is.

Sources & References

Michael C.

Michael C.

Founder & Principal Engineer, Selina Labs

Michael builds Selina, a privacy-first AI that remembers you across conversations. He ships security-sensitive AI in production — real attacks, real fixes, measured in minutes and dollars — and writes about privacy, security, and LLMs from that seat. Top Rated Plus and expert-verified on Upwork.

Learn more about Selina.ai