SELINA.ai
Sign in

Is ChatGPT Confidential? What Actually Happens to the Data You Paste In

You're about to paste something sensitive into ChatGPT. Maybe it's a contract clause, a patient summary with the names swapped out, a chunk of proprietary code, or a personal journal entry you want help rewriting. Before you hit enter, you want to know: is ChatGPT confidential? The short answer is no, not in any sense a lawyer, a compliance officer, or a reasonably paranoid engineer would recognize. The longer answer depends on which product tier you're using, which settings you've toggled, and whether a federal court decides it wants to see your logs.

Key Takeaways

What Does OpenAI Actually Do with Your Input?

It depends on your account type. On consumer accounts (free and Plus), OpenAI's policy states that your conversations may be used to improve model performance unless you explicitly opt out. That opt-out lives in Settings → Data Controls, or through the privacy portal. Most users never touch it.

When you opt out, your conversations stop flowing into the training pipeline. They do not, however, vanish. OpenAI retains conversations for up to 30 days even after you delete them, for the stated purpose of abuse and safety monitoring. So the data sits on their servers, accessible to their trust-and-safety team, for a month after you thought it was gone.

For business tiers, the picture changes. OpenAI states that ChatGPT Team, Enterprise, and API inputs are not used for training by default. Deleted conversations are removed from their systems within 30 days unless a legal obligation requires otherwise. The "unless" clause is doing a lot of work in that sentence, and we'll get to why.

Does Opting Out of Training Make Your Data Private?

No. It makes your data not-trained-on. Those are different things.

When you flip the training toggle off, you've addressed one specific use of your data. Your input is still transmitted to OpenAI's servers, processed by their models, and stored temporarily. OpenAI employees with appropriate access can still review it for safety purposes. The content still exists within a system you do not own, operate, or audit.

Think of it this way: telling a hotel you don't want housekeeping to enter your room doesn't mean the hotel doesn't have a master key. Opting out of training removes one specific use. It does not change the underlying access architecture.

Can a Court Force OpenAI to Hand Over Your Conversations?

Yes. This has already happened.

In June 2025, a federal court overseeing the New York Times copyright lawsuit against OpenAI ordered the company to preserve all ChatGPT logs, including chats that users had deleted and conversations from temporary (non-logged-in) sessions that would normally have been purged under the 30-day policy. OpenAI objected. The court didn't care.

That blanket preservation order was eventually narrowed and lifted in September 2025, but not before months of user data that should have been deleted under normal policy was saved under legal hold. And in November 2025, a federal judge affirmed an order requiring OpenAI to produce 20 million anonymized ChatGPT logs to the plaintiff news organizations.

The implications here are structural, not incidental. Any company that stores your data, even temporarily, is subject to litigation holds and discovery obligations. A deletion policy is a promise about what the company intends to do with your data in the normal course of business. A subpoena overrides that promise. If your data exists on someone else's servers when the subpoena lands, it gets preserved, regardless of what the privacy policy says.

This is not a hypothetical risk. It's a documented event that affected real user data.

What Happened with the Google Indexing Incident?

In 2025, OpenAI offered a sharing feature that let users make ChatGPT conversations "discoverable" by search engines. Roughly 4,500 shared conversations ended up indexed by Google. No private, unshared conversations were exposed, but the shared conversations that were indexed contained material users probably didn't expect to be publicly searchable.

Some of those conversations included personal names, email addresses, résumés, children's names, and emotionally sensitive disclosures. People had used the share function to send a conversation to a colleague or a friend, not realizing that "discoverable" meant "findable by anyone with a search engine."

OpenAI pulled the discoverable sharing option in August 2025 and worked to de-index previously exposed links. But the episode illustrates something important: the most likely vector for your data leaking out of ChatGPT is not model training. It's product surface area. Features, defaults, integrations, sharing mechanisms. The stuff that gets bolted on around the edges.

Separately, around the same time, unusually long user queries started appearing inside Google Search Console reports, apparently the result of an indexing or routing bug rather than intentional scraping. Another reminder that data flowing through complex systems finds unexpected exits.

Does Attorney-Client Privilege Survive if You Use ChatGPT?

Almost certainly not, based on the rulings so far.

A federal judge ruled that submitting information to a public AI platform counts as disclosure to a third party, which waives attorney-client privilege and confidentiality protections. The reasoning is straightforward: when you send data to a commercial AI service, you've shared it with a company. That company stores it, processes it, and may have employees who can access it. You've broken the circle of confidentiality.

A related ruling in a criminal case, the Heppner decision in early 2026, reached the same conclusion through slightly different reasoning: an AI is not a lawyer, and the vendor's standard privacy policy negates any reasonable expectation of confidentiality. You can't claim privilege over a communication that you voluntarily sent to a commercial third party whose terms of service give it broad rights to process, store, and potentially review that content.

This is new law. More rulings will come. But the direction of travel is clear. If you're a lawyer and you paste privileged client information into a consumer AI tool, you're taking a real risk that you've just waived the privilege. The same logic extends, less formally, to any professional confidentiality obligation: medical, financial, fiduciary.

How Do Enterprise and API Tiers Differ?

OpenAI's Enterprise privacy page states that for Enterprise, Business, and Edu customers, the company only receives rights in input and output necessary to provide the service, comply with law, and enforce policies. No training on your data. Deleted conversations removed within 30 days.

These are real, meaningful differences from the consumer tier. They are also contractual differences, not architectural ones. Your data still transits OpenAI's infrastructure. It still sits on their servers during processing. It is still subject to the same legal process risks (subpoenas, litigation holds) as any data stored by any third party.

The Enterprise tier gives you a better contract. It does not give you a different physics. If the data exists on someone else's computer, someone else can be compelled to produce it.

What Does "Confidential" Actually Mean Here?

When people ask "is ChatGPT confidential," they usually mean one of three things, and the answer is different for each.

"Will OpenAI employees read my conversations?" Probably not routinely. But they can, for safety review and abuse monitoring. The policy permits it. The 30-day retention window exists specifically to enable it.

"Will my data be used to train models?" On consumer accounts, yes by default, no if you opt out. On Enterprise/API, no by default. This is the question most coverage focuses on, and it's actually the least important one from a confidentiality standpoint.

"Is my data protected from disclosure to third parties?" Not in any robust sense. It can be subpoenaed. It can be subject to litigation holds. It has, in at least one case, been indexed by search engines through a product feature. Courts have ruled it does not carry privilege. "Confidential" in the way a lawyer or a doctor uses that word? No.

Why Training Is the Wrong Thing to Worry About

Most of the public conversation about ChatGPT and privacy centers on training data. Will your conversations end up baked into the model? It's a reasonable concern, but it's not where the real exposure is.

The actual risks that have materialized so far are all about storage and access, not training. Courts compelling log production. Sharing features exposing conversations to search engines. The 30-day retention window creating a surface for subpoena. These are infrastructure and policy risks, not machine-learning risks.

A model trained on your data is a diffuse, probabilistic risk. It might, in theory, regurgitate something resembling your input in response to a specific prompt. The probability is low and the output would be noisy. A court ordering the production of your actual, verbatim conversation logs is a concrete, deterministic risk. Your exact words, in a legal filing, attributed to your account.

Focus on storage. Focus on retention. Focus on who can be compelled to hand over what.

What Would Actual Confidentiality Look Like?

If you want your AI interactions to be genuinely confidential, the architecture has to enforce it, not just the policy. A few properties matter.

First: the data should not be stored in a form that the vendor can read after processing is complete. If the vendor can read it, the vendor can be compelled to produce it. If the vendor can't read it, there's nothing meaningful to produce in response to a subpoena.

Second: files and data transfers should be end-to-end encrypted so that the vendor never has access to the plaintext. This is technically achievable for file transfer and storage. It is harder for the parts of a system that need to process your input through a language model, because the model needs to see the input to generate a response.

Third: retention should be minimal by design. Not "we promise to delete it in 30 days" but "we architecturally cannot retain it beyond what's needed for the immediate request." A short retention window for operational metadata is realistic. Storing full conversation logs for a month is a choice, and it's a choice that creates exposure.

This is the design philosophy behind Selina. Content is encrypted at rest. Files and transfers through SelinaSEND are zero-knowledge encrypted. Memory is not end-to-end encrypted, because a slice of each request reaches a frontier provider at inference, and we think it's important to say that plainly rather than imply otherwise. Non-content operational metadata is kept for a short retention window, not indefinitely. The account itself is protected, not encrypted, because there's a difference and the difference matters.

None of this makes any system perfectly confidential. But there's a structural gap between "we contractually promise not to look at your data" and "we architecturally cannot access your data at rest." The former is a policy. The latter is a constraint. Policies can be overridden by courts, by acquisitions, by policy changes, by breaches. Constraints hold unless the math breaks.

What Should You Do Before Pasting Sensitive Data into Any AI Tool?

Ask four questions.

  1. What tier am I on, and what does the contract actually say? Consumer terms and enterprise terms are different products with different data-handling obligations. Read the specific terms for your specific tier. The defaults differ significantly between free, Plus, Team, Enterprise, and API access.
  2. What happens to this data if the vendor gets sued? Not "what does the vendor promise" but "what can a court compel." If the vendor stores your data for any period, it's discoverable. If the vendor has already been subject to one major litigation hold (and OpenAI has), you should assume it could happen again.
  3. Would disclosure of this specific input create a concrete harm? Pasting a draft blog post is different from pasting a client's medical records. Calibrate your behavior to the sensitivity of the content, not to a general feeling about privacy.
  4. Does this input contain information subject to a professional or legal confidentiality obligation? If you're a lawyer, a doctor, a therapist, or anyone holding information under a duty of confidentiality, the answer to "is this AI tool confidential enough" is almost certainly no for consumer-tier products. The Heppner decision and related rulings make this clear.

The Difference Between a Policy and an Architecture

OpenAI's privacy policies are, as corporate privacy policies go, relatively clear. They tell you what they do with your data, tier by tier. A 2026 independent review gave OpenAI a privacy score of 48 out of 100, a C grade, noting the gap between what users assume and what the policies actually permit.

The problem is not that the policies are deceptive. The problem is that policies are words in a document. They describe intentions, not constraints. They can be changed with 30 days' notice. They can be overridden by legal process. They are enforced by trust, not by cryptography.

When someone asks "is ChatGPT confidential," they're usually asking whether they can trust the system with sensitive information. The honest answer is: you can trust it about as much as you can trust any third-party cloud service that stores your data in readable form, processes it on infrastructure you don't control, and operates in a jurisdiction where courts can compel disclosure. Which is to say: it depends on your threat model, your professional obligations, and how much you'd mind if the input showed up in a legal filing three years from now.

For casual use, for brainstorming, for drafting content that isn't sensitive, ChatGPT is fine. For anything where confidentiality is a legal or ethical requirement, the architecture matters more than the policy. Look for systems where the vendor cannot access your data at rest, where retention is minimal by design rather than by promise, and where encryption is structural rather than contractual.

If you want an AI assistant built around those constraints, start a free 7-day trial of Selina, no card required.

Frequently Asked Questions

Is ChatGPT actually confidential?

No, not in the sense a lawyer or compliance officer would recognize. Protections vary by account tier and settings, but data still passes through infrastructure you don't control and can be subject to legal process.

Does opting out of training mean my ChatGPT data is deleted?

No, opting out only stops your conversations from being used for training; OpenAI still stores conversations for up to 30 days for abuse and safety monitoring, and staff with appropriate access can review them.

Can courts force OpenAI to hand over ChatGPT conversations?

Yes, this has already happened; a federal court ordered OpenAI to preserve all ChatGPT logs, including deleted and temporary chats, in the New York Times copyright lawsuit, and later ordered production of 20 million anonymized logs.

Does using ChatGPT put attorney-client privilege at risk?

Almost certainly yes, as federal rulings have held that submitting information to a public AI platform counts as disclosure to a third party, which can waive attorney-client privilege and other confidentiality protections.

Are Enterprise and API tiers more private than the free version of ChatGPT?

They offer stronger contractual protections, such as no training on data by default and shorter retention, but the underlying architecture is the same, data still transits and sits on OpenAI's servers and remains subject to subpoenas and litigation holds.

Sources & References

Michael C.

Michael C.

Founder & Principal Engineer, Selina Labs

Michael builds Selina, a privacy-first AI that remembers you across conversations. He ships security-sensitive AI in production — real attacks, real fixes, measured in minutes and dollars — and writes about privacy, security, and LLMs from that seat. Top Rated Plus and expert-verified on Upwork.

Learn more about Selina.ai