California AB 2013: Training Data Transparency for GenAI Developers
California AB 2013 takes effect 1 January 2026 and requires generative AI developers to publish summaries of their training data. Here is who it covers and what you have to disclose.
California AB 2013 is a training-data transparency law that takes effect on 1 January 2026. It is narrow but sharp: if you develop generative AI systems, you have to tell the public what you trained them on.
It matters far beyond California. The state hosts a large share of the world's AI developers, and disclosure obligations tend to become de-facto global standards — once a summary is published for one market, it is published for everyone.
Who this actually covers
AB 2013 applies to developers of generative AI — the organisations building or substantially modifying the models themselves.
- If you build or fine-tune a generative model, you are in scope.
- If you only use third-party AI tools, you are not the developer and this law does not bind you directly.
That distinction matters, because most SMEs sit firmly in the second group. The realistic impact for a small business is not compliance — it is leverage. Your vendors will soon be publishing what their models were trained on, and that is information you can use.
What developers must publish
The core obligation is a public summary of the training dataset covering:
- Sources of the data used to train the system.
- Licensing and ownership status of that data.
- Whether the data includes personal information.
- Whether the data includes synthetic data.
This is a disclosure requirement, not a data-sharing one. You are describing the dataset, not handing it over.
The interesting consequence of AB 2013 is not what it forces developers to say. It is that, for the first time, buyers can compare vendors on the provenance of their training data.
What this means for an SME
If you do not build models, treat AB 2013 as a procurement tool rather than a compliance burden:
- When evaluating an AI vendor, ask for their training-data summary. From 2026, serious GenAI developers will have one.
- Check whether the summary mentions personal information. If it does, that has direct implications for your own GDPR or state-privacy position when you feed customer data into that tool.
- Watch for vendors who cannot produce a summary at all. That is a signal about their maturity, not just their legal exposure.
If you do build or fine-tune models — including fine-tuning an open model on your own data — you may be closer to developer status than you assume. Fine-tuning is exactly the kind of substantial modification that can pull a business into scope.
Practical next steps
- Confirm honestly whether you develop, fine-tune, or merely use generative AI. The answer decides everything else.
- If you fine-tune, start documenting your data sources now. Retrofitting provenance records after the fact is painful and often impossible.
- If you only use AI tools, add "training data summary" to your vendor due-diligence checklist.
- Keep this alongside your wider US position — California is one law in a fast-growing patchwork that already includes Colorado and Texas.
Frequently asked questions
Does AB 2013 apply if I only use GenAI tools rather than build them?
No. AB 2013 targets developers of generative AI systems. If you only use third-party tools such as ChatGPT or Copilot, this law does not create obligations for you — though your vendor's disclosures become a useful due-diligence source.
What counts as a training data summary?
A published summary describing the sources of the data, its licensing or ownership status, and whether it includes personal information or synthetic data. It is a disclosure obligation, not a requirement to publish the dataset itself.