SELINA.ai
Sign in

The 2026 State Privacy Law Patchwork Just Got a New Layer: Data Broker Registration and AI Training Data Rules

If you ship a product that touches consumer data in the United States, the privacy compliance surface just expanded again. New Jersey signed a data broker law with a $1.5 million top-tier registration fee. Connecticut now requires you to disclose whether you use personal data to train AI models. California wants you to publish your training data sources and update the disclosure every time you retrain. And the federal government is making noises about preemption while doing nothing concrete to preempt. Here is what actually matters for your product, right now, in mid-2026.

Key Takeaways

Which States Actually Have Data Broker Registration in Force Right Now?

Four: California, Vermont, Texas, and Oregon. These are the states where registration is mandatory and operational today. Several others, including New Jersey, Delaware, Michigan, and Alaska, are at various stages of building out their frameworks, but the registries either don't exist yet or aren't accepting filings. If you are prioritizing compliance work, those four are where real enforcement risk sits in July 2026.

That count is about to grow, but the timing is messy. New Jersey's law is technically effective immediately. The registry to actually register with does not exist. More on that below.

What Does New Jersey's A5328 Actually Do?

Governor Mikie Sherrill signed A5328 on June 30, 2026, making New Jersey the seventh state with a data broker law. The headline number is the $1.5 million annual registration fee for the largest brokers, structured in tiers based on volume. That fee is eye-catching. It is not the part that should keep you up at night.

The part that matters: New Jersey is the first state to regulate "data collectors" alongside traditional data brokers. Most state broker laws only cover companies with no direct relationship to the consumer. If you collect an email at checkout and then sell or license that data downstream, existing laws in other states generally don't treat you as a broker. New Jersey closed that gap. A retailer, a SaaS platform, an app with a login screen: if you have a direct consumer relationship and you license data to a third party, you're now a "data collector" under A5328, and the law applies to you.

The practical self-test is short. Do you sell or license any customer data, even incidentally via an integration partner, an analytics SDK, or an ad pixel? If yes, and if New Jersey residents are in your user base, you should be tracking A5328. Company size and intent are not the trigger. The data flow is.

Is New Jersey's Law Actually Being Enforced?

Not yet. This is one of the stranger situations in the current patchwork. The law's obligations are technically effective immediately. But the Division of Consumer Affairs issued an alert on July 10, 2026 stating that brokers won't need to register or pay fees until the registry launches, which is projected for spring 2027. A senior state official reportedly indicated the governor won't enforce until the legislature addresses certain "defects" in the statute.

So you have a law on the books, no registry to register with, and a public signal that enforcement is on hold. This pattern is becoming common. It creates a specific kind of compliance-fatigue problem: you can't just track effective dates, you need to track enforcement readiness. The legal exposure is latent, and when the registry does launch, there will likely be a scramble.

Our read: if you would be covered under the data-collector provision, start building toward compliance now. The spring 2027 window will arrive faster than your legal team's backlog will clear.

What Changed in California's Broker Rules for 2026?

California's SB 361, the Defending Californians' Act, went into effect in January 2026. It expanded disclosure requirements for registered brokers. The new obligations include revealing whether data includes sensitive categories like citizenship or immigration status, and whether it has been shared with foreign actors, government or law enforcement agencies, or generative AI developers.

That last category is the one with teeth for AI founders. If you are a registered broker in California and you provide data to any generative AI developer (including yourself, if you are both), you now have to disclose that fact. The disclosure is not optional and not buried in a general privacy policy. It is a specific line item in your broker registration.

Separately, California's Delete Request and Opt-Out Platform (DROP) went live for consumers on January 1, 2026. DROP lets California residents submit a single deletion request that propagates to all registered brokers. Starting August 1, 2026, brokers must check DROP at least every 45 days and process deletions within 90 days. If you are a registered broker in California, you need a technical integration with DROP or a manual process that meets that cadence. Forty-five days is not a lot of runway if your data pipeline doesn't support granular deletion.

Does Connecticut's New Disclosure Rule Apply to My Product?

Probably, if you use consumer data anywhere in an AI training pipeline and you have Connecticut residents in your user base. Connecticut's amended Data Privacy Act took effect July 1, 2026, and it requires businesses to disclose in their privacy notice whether they use personal data to train AI models. This makes Connecticut the first state to require AI-training transparency as a standalone privacy-notice item.

The important detail: Connecticut doesn't define "large language model." Compliance lawyers are reading this broadly to cover any use of consumer data to fine-tune, adapt, or pre-train a foundation model. If you fine-tune a model on customer support tickets, that counts. If you use interaction logs to improve a recommendation system built on a transformer architecture, that likely counts too. The threshold is not "are you building a large language model from scratch." The threshold is "do you use personal data in any model training process."

And you have to state it even if the answer is "no." The law requires affirmative disclosure either way. If you don't train on consumer data, say so in your notice. If you do, say so and describe it. Silence is not compliant.

What Does California's AB 2013 Require for AI Training Data Transparency?

AB 2013, effective January 1, 2026, requires developers of generative AI systems to publicly post details on their training data. The required disclosures include: sources of training data, types of data used, copyright status, and whether personal information is included. This information must be updated whenever the system undergoes a "substantial modification."

Here is where it gets operationally painful. A "substantial modification" under AB 2013 includes any new version or update that materially changes functionality or performance, and that explicitly includes changes from retraining or fine-tuning. If you run a continuous fine-tuning pipeline (which many production AI systems do), every meaningful retrain could trigger a disclosure update obligation.

This is a product design constraint, not just a legal checkbox. If your architecture fine-tunes on rolling windows of user data, you are creating a rolling compliance obligation. Every sprint that ships a retrained model is potentially a sprint that also requires a disclosure update. The legal and engineering cadences are now coupled.

How Does This Affect Founders Who Don't Train on User Data?

You are in a structurally better position, and that gap is widening. If your architecture never ingests consumer data into a training pipeline, Connecticut's disclosure is a one-line addition to your privacy notice ("we do not use your personal data to train AI models") and AB 2013's ongoing update obligation doesn't compound with your release cycle.

At Selina, we don't train on your conversations. Memory is encrypted at rest (though memory is NOT end-to-end encrypted, since a slice of each request reaches a frontier provider at inference). Files and transfers via SelinaSEND are end-to-end encrypted. We route requests through a stack of frontier models, routed per task, and none of your content feeds back into model training. Non-content operational metadata is kept for a short retention window.

This isn't just an ethical stance. It's becoming the only way to avoid a moving compliance target that shifts every time you push a model update. Privacy-by-design architectures, where you simply don't train on customer data, are turning into a competitive advantage measured in hours of legal review saved per quarter.

What Happened with Colorado's AI Act?

Colorado provides a useful case study in regulatory instability. The state repealed its original AI Act on May 14, 2026 and replaced it with SB 189, a narrower statute focused on automated decision-making technology used in "consequential decisions." The replacement takes effect January 1, 2027, unless delayed by pending litigation.

The backstory matters. In April 2026, a lawsuit argued the original law was unconstitutionally vague and violated the First Amendment, with the Department of Justice intervening on the plaintiff's side. That intervention was notable: the first time the federal government sought to invalidate a state AI law. The legislature responded by repealing and replacing rather than defending the original text in court.

The lesson for founders: a law's effective date is not the same as its durability. Colorado's original AI Act never actually took effect in its original form. If you had built compliance infrastructure specifically for that statute's requirements, some of that work is now stranded. The replacement has different scope, different definitions, and different obligations.

Will Federal Preemption Make State Laws Irrelevant?

Not anytime soon. Executive Order 14365, issued December 11, 2025, signals an intent to limit states' authority to enact and enforce individual AI laws. The order carves out exemptions for child safety, AI compute and data-center infrastructure, and state government use of AI. But an executive order is not legislation. Congress has not passed a preempting federal law.

The prudent approach is to keep complying with state laws until federal preemption is enacted by statute, not just signaled by executive order. The current EO could be reversed by a future administration, and in the meantime, state attorneys general retain enforcement authority over their own statutes.

The scale of the patchwork gives you a sense of why preemption keeps coming up. More than 35 states had active AI legislation as of March 2026. Across 48 states, 478 bills are being tracked covering watermarking, chatbot disclosure, synthetic media labeling, and training-data transparency, with 40 bills taking effect in 2026 alone. That is the environment you are building in.

How Many State Privacy Laws Are in Effect Right Now?

Twenty states now have comprehensive privacy laws in effect, with Indiana, Kentucky, and Rhode Island joining in 2026. Effective dates are staggered across January 1, July 1, and August 1. This doesn't count the data broker registration laws or AI-specific statutes, which layer on top.

For a founder with a nationally available product, the compliance matrix is roughly: 20 comprehensive privacy laws, 4 operational broker registries (with more coming), at least 2 states with AI training data disclosure requirements (Connecticut and California), and an uncountable number of pending bills that could change the landscape in any legislative session. This is not a problem you solve once. It is a recurring cost of doing business.

What Should You Actually Do Right Now?

Run through this list. It is not exhaustive, but it covers the high-priority items from the 2026 changes.

  1. Audit your data flows for the New Jersey "data collector" trigger. Do you license, sell, or share customer data with any third party, even through an SDK, ad pixel, or analytics integration? If yes, and if you have New Jersey users, start preparing for A5328 compliance before the registry launches in spring 2027.
  2. Update your privacy notice for Connecticut. Add an affirmative statement about whether you use personal data to train AI models. This is required as of July 1, 2026, regardless of whether the answer is yes or no.
  3. If you develop generative AI and operate in California, publish your AB 2013 disclosures. Training data sources, data types, copyright status, personal information inclusion. Build a process for updating these disclosures on every substantial modification, including retrains and fine-tunes.
  4. If you are a registered broker in California, integrate with DROP. The 45-day check cadence starts August 1, 2026. If your data architecture doesn't support granular, consumer-level deletion, that is a technical problem you need to solve before the deadline.
  5. Track enforcement readiness, not just effective dates. New Jersey's law is effective but unenforceable. Colorado's AI Act was rewritten before it took effect. The gap between "law on the books" and "law being enforced" is where compliance-fatigue waste accumulates. Build your tracking around enforcement milestones (registry launch dates, AG enforcement actions, rulemaking deadlines) rather than just statutory effective dates.
  6. Consider whether your architecture avoids the problem entirely. If you don't train on consumer data, the AI training disclosure rules are a one-time notice update rather than an ongoing engineering burden. That architectural choice has a measurable compliance cost difference.

Why the Enforcement Gap Matters More Than the Law Count

The pattern across 2026 is consistent: legislatures pass laws faster than regulatory agencies can build the infrastructure to enforce them. New Jersey's registry won't exist for almost a year after the law's effective date. Colorado's original AI Act was repealed before enforcement began. Federal preemption is a policy intention without a statutory vehicle.

This creates a specific kind of risk for founders. You can't ignore laws just because enforcement hasn't started. The statutes exist, and a future enforcement action could be retroactive to the effective date. But you also can't treat every bill as equally urgent, or you'll spend all of your compliance budget on laws that may never take effect in their current form.

The practical approach: prioritize by enforcement readiness and your actual data flows. A law in a state where you have users, where the registry is live, and where your product handles covered data is a real compliance task. A law in a state where you have no users and the enforcement infrastructure doesn't exist yet is a monitoring task. Know which is which.

The Architectural Argument for Not Training on User Data

We keep coming back to this because the regulatory environment is making it more true every quarter. Every new training-data disclosure law adds ongoing cost to architectures that feed user data into model training. Connecticut's disclosure. California's AB 2013 update-on-every-retrain requirement. California's SB 361 broker disclosure about sharing data with AI developers. These stack.

If you design your system so that user content never enters a training pipeline, you neutralize entire categories of compliance work. You still need privacy notices. You still need to handle deletion requests. But you don't need to update your training data disclosure every time you push a model version, because there's nothing to disclose.

We built Selina this way. Your conversations aren't training data. Your files transferred via SelinaSEND are zero-knowledge encrypted. Your account is protected, your content is encrypted. We are honest about the limits: memory is NOT zero-knowledge encrypted, because inference requires sending data to a frontier provider. But the training-data question has a clean answer: we don't do it.

That clean answer saves us, conservatively, dozens of hours per quarter in legal review and disclosure updates. It will save more as additional states adopt similar rules. And it lets us tell you, without hedging, exactly what happens to your data.

If you want an AI assistant built on that principle, start a free 7-day trial, no card required.

Frequently Asked Questions

Which states currently have operational data broker registration requirements?

California, Vermont, Texas, and Oregon have mandatory, operational registries today. Other states like New Jersey, Delaware, Michigan, and Alaska have laws in progress but no functioning registry yet.

Does New Jersey's A5328 apply to companies that aren't traditional data brokers?

Yes. A5328 creates a new 'data collector' category covering companies with direct consumer relationships that sell or license customer data, even incidentally through integrations, analytics SDKs, or ad pixels. This means retailers, SaaS platforms, or apps with logins can be covered even though they aren't traditional brokers.

Is New Jersey actually enforcing its data broker law right now?

No. Although the law is technically effective immediately, the state's registry won't launch until spring 2027, and the Division of Consumer Affairs has said brokers don't need to register or pay fees until then. A senior official also indicated the governor won't enforce until the legislature fixes certain statutory defects.

What does Connecticut's new AI disclosure requirement actually require?

Effective July 1, 2026, Connecticut requires businesses to state in their privacy notice whether they use personal data to train AI models, and the disclosure is mandatory even if the answer is no. The definition isn't limited to large language models and is being read broadly to include fine-tuning or adapting any model type on consumer data.

How often do companies need to update AI training data disclosures under California's AB 2013?

AB 2013 requires generative AI developers to update their public training data disclosures whenever there's a 'substantial modification,' which explicitly includes retraining or fine-tuning. For companies running continuous fine-tuning pipelines, this can mean an update obligation tied to nearly every retrain cycle.

Sources & References

Michael C.

Michael C.

Founder & Principal Engineer, Selina Labs

Michael builds Selina, a privacy-first AI that remembers you across conversations. He ships security-sensitive AI in production — real attacks, real fixes, measured in minutes and dollars — and writes about privacy, security, and LLMs from that seat. Top Rated Plus and expert-verified on Upwork.

Learn more about Selina.ai