California enacted an “Artificial Intelligence Training Data Transparency” statute. Cal. Civ. Code §3111. It “requires developers of ‘a generative artificial intelligence system or service’ that is ‘publicly available to Californians for use’ to ‘post on the developer’s internet website documentation regarding the data used by the developer to train the generative artificial intelligence system or service.’” X.AI LLC v. Bonta, 2026 WL 626926 (C.D. Cal. Mar. 4, 2026).
Plaintiff, X.AI, produces and develops A.I. models that it shares with the public. It filed suit to enjoin enforcement and moved for a preliminary injunction. The motion was denied. However, the court left the door open and wrote that the preliminary injunction decision was only a “threshold inquiry.”
THE SCOPE OF A.I. TRAINING IS IMPORTANT
The X.AI court’s analysis is of interest to general civil and criminal litigation because the scope of a GenAI training set may be important to laying, or challenging, an evidentiary foundation to a proffer of A.I. generated evidence at trial or on motions. For example, if a training set is biased, the output may be biased. See A Review of Sedona’s “Artificial Intelligence (AI) and the Practice of Law” by The Hon. Xavier Rodriguez (Sep. 27, 2023)(Sedona suggests: “AI evidence may require that the offering party disclose any training data used by the AI platform to generate the exhibit.”)(emphasis added).
THE CALIFORNIA TRAINING DATA TRANSPARENCY ACT
The X.AI court described the California statute:
The documentation must include “[a] high-level summary of the datasets used in the development of the generative artificial intelligence system or service” addressing, but not limited to, twelve enumerated topics…. Those topics include:
(1) The sources or owners of the datasets.
(2) A description of how the datasets further the intended purpose of the artificial intelligence system or service.
(3) The number of data points included in the datasets, which may be in general ranges, and with estimated figures for dynamic datasets.
(4) A description of the types of data points within the datasets….
(5) Whether the datasets include any data protected by copyright, trademark, or patent, or whether the datasets are entirely in the public domain.
(6) Whether the datasets were purchased or licensed by the developer.
(7) Whether the datasets include personal information ….
(8) Whether the datasets include aggregate consumer information ….
(9) Whether there was any cleaning, processing, or other modification to the datasets by the developer, including the intended purpose of those efforts in relation to the artificial intelligence system or service.
(10) The time period during which the data in the datasets were collected, including a notice if the data collection is ongoing.
(11) The dates the datasets were first used during the development of the artificial intelligence system or service.
(12) Whether the generative artificial intelligence system or service used or continuously uses synthetic data generation in its development….
Id. at *1-2. However:
The statute exempts three types of models from disclosures: (1) a generative artificial intelligence system or service whose sole purpose is to help ensure security and integrity; (2) a generative artificial intelligence system or service whose sole purpose is the operation of aircraft in the national airspace; and (3) a generative artificial intelligence system or service developed for national security, military, or defense purposes that is made available only to a federal entity.
Id. at *2.
X.AI’s CLAIMS
X.AI is an entity subject to the statute. It presented three arguments against the statute: “(1) that it violates the Takings Clause of the Fifth Amendment; (2) that it violates the First Amendment; and (3) that it is unconstitutionally vague.”
Presumably as a precaution, X.AI published a “high-level, limited disclosure that does not reveal its trade secrets.” Id. at *2. In is Complaint, however, it alleged a concern that the State Attorney General would assert non-compliance and seek to enforce the law.
Plaintiff alleges that several aspects of the datasets used to train AI models—including their contents, origins, size, and cleaning methods—are valuable and non-public…. Plaintiff alleges that “information about the datasets and processes AI developers use to train their AI models is a closely protected trade secret.”
Id. at *2. The X.AI court described and applied the legal standard governing motions for preliminary injunction.
THE COURT’S DECISIONS
The Attorney General of California was the defendant. I will, for ease of reference, even if not with fidelity to U.S. Const., Amend. 11, refer to the defendant as the “State.”
X.AI Has “Standing” to Challenge the Statute
The State contended that X.AI lacked Art. III “standing” to make the claim. Oversimplifying, the State argued that X.AI had “engaged in some degree of compliance” and lacked injury in fact. The court rejected that defense. Id. at *3.
The Statute Did Not Violate the Takings Clause Because the Complaint was Generalized
The X.AI court next analyzed the “Takings Clause” argument.
Before the Court can evaluate whether Plaintiff has likely allegedly a successful Takings Clause claim, the Court must determine the likelihood of Plaintiff proving that the sources, sizes, and cleaning methods of its datasets qualify as trade secrets…. Under California law, “the test for a trade secret is whether the matter sought to be protected is information (1) that is valuable because it is unknown to others and (2) that the owner has attempted to keep secret.” [emphasis added].
However, the X.AI court found that X.AI’s supporting allegations were simply general statements. It wrote: “ Plaintiff’s Complaint trades in frequent abstraction and hypotheticals, rather than pleading specifics about Plaintiff’s practices.” Id. at *4. The court explained that:
When it comes to specificity, Plaintiff alleges that “[a]s part of its development process, xAI generally used the methodology outlined above.” … It offers that “xAI’s engineers invested substantial amounts of time and energy in acquiring datasets from various sources across the Internet to develop and eventually train the AI models that it has produced.” … Plaintiff also alleges that “[t]he amount of data that xAI uses is also valuable precisely because it is unknown to others” and that “xAI’s processes for cleaning, modifying, and refining the datasets it has obtained are economically valuable information too.”
Id.
Further, X.AI acknowledged that there is data overlap among many AI companies. Instead, it asserted that the differences give a competitive edge. The court wrote: “The problem is that Plaintiff has not alleged that it actually uses datasets that are unique, that it has meaningfully larger or smaller datasets than competitors, or that it cleans its datasets in unique ways. Plaintiff’s resort to generalizations and hypotheticals about the AI model development industry make it difficult for the Court to find that Plaintiff has carried the heavy burden of showing a likelihood of success in proving that trade secrets are at play here.” Id. at *4.
The X.AI court acknowledged that, hypothetically, datasets could be trade secrets; however, here, X.AI’s allegations were only an “abstract pleading….” Id. at *5. It held that: “Plaintiff has failed at this stage to sufficiently allege that trade secrets are implicated. As such, the Court finds that, as a threshold matter, Plaintiff is not likely to succeed on the merits of its Takings Clause claim based on the Complaint as pled.”
The Statute Did Not Violate the First Amendment
X.AI asserted that the statute compelled speech based on content and viewpoint. “Specifically, Plaintiff alleges that [the California statute] is content-based because it requires Plaintiff to disclose specific content about its AI models, and it is viewpoint-based because it exempts developers of AI models related to network security, aircraft operations, and national security from its requirements.” Id. at *5.
The X.AI court found that the statute is a “content-based speech regulation….” Id. at *6. However, it also found it to be “commercial speech.” Id. at *8. It wrote that: “No part of the statute indicates any plan to regulate or censor models based on the datasets with which they are developed and trained.” Id. at *7.
After a lengthy analysis, the X.AI court concluded:
Ultimately, Plaintiff has demonstrated a distinct possibility of prevailing on the merits…. But it had not demonstrated a likelihood of success on the merits. The information before the Court is insufficient to come to such a conclusion at this stage. Plaintiff therefore does not satisfy this threshold inquiry for a preliminary injunction on its First Amendment claim.
Id. at *8 (emphasis in original).
The California Statute is Not Void for Vagueness
The X.AI court wrote that the mandate of publishing “a high level summary: of the training datasets is not a picture of clarity standing alone….” Id. at *9. However, that mandate is followed “by a precise list of the information to be included.” Id.
X.AI challenged terms such as “dataset” and “data point” as vague. It also argued that the list was non-exhaustive and therefore vague. The court disagreed.
Here, there is a list of information required akin to a set of factors—it is simply non-exhaustive. Given that a statute entirely lacking a list of factors can still be sufficiently clear, it is likely that a non-exhaustive list is enough.
Id. at *9.
The court viewed other vagueness challenges as “similarly insufficiently persuasive at this stage, absent a better-developed record, to find a likelihood of success on the merits…. Ultimately, the record at this stage is insufficiently developed for the Court to determine that Plaintiff is likely to succeed on the merits of its vagueness challenge.” Id.
However, the X.AI court left the door open: “Evidence may arise during the course of litigation that eventually requires a different determination. But the pleadings and record as they stand are not enough at this time.” Id.
COMMENT
Knowledge of how an A.I. tool was trained may be important in offering, or challenging, evidence at trial. Admissibility may turn, at least in part, on knowing the data on which the AI was trained. See M. Grossman & Hon. P. Grimm, “Judicial Approaches to Acknowledged and Unacknowledged AI-Generated Evidence,” 26 Col. Sci. & Tech. L. Rev. 110, 152 (2025).
ABA Formal Opinion 512 (“Generative Artificial Intelligence Tools” 2024) states: “The large language models underlying GAI tools use complex algorithms to create fluent text, yet GAI tools are only as good as their data and related infrastructure. If the quality, breadth, and sources of the underlying data on which a GAI tool is trained are limited or outdated or reflect biased content, the tool might produce unreliable, incomplete, or discriminatory results.” [emphasis added].
GenAI responds to prompts “based on patterns and structures learned from the data used to train the AI model.” Maryland State Bar Ass’n., “An Overview of Ethical Considerations for Attorney Use of Generative Artificial Intelligence Technologies” (undated), 3 (emphasis added).
The “quality of their training” may impact the limitations and risks presented by GenAI tools. The Sedona Conference, The Sedona Canada Primer on Artificial Intelligence and the Practice of Law, 26 SEDONA CONF. J. 103,128 (forthcoming 2025), 128. A comprehensive and representative dataset is needed to train AI systems. Id. at146-47. Inadequate training may lead to bias. Id. at 149-50.
Disclosure of training data may be an important predicate to admissibility. American Assoc. for the Advancement of Science, “Artificial Intelligence and the Courts” (2022), 12.
Facial recognition technology is a form of artificial intelligence. In criminal cases, Maryland’s facial recognition technology statute requires a “description and the names of the databases searched….” Md. Code Ann., Crim. Proc. Art. §2-504.
The ABA Task Force on Law and Artificial Intelligence Year 2 Report (2025) states: “Similarly, in healthcare, AI-trained devices may produce biased results if training datasets lack diversity, potentially leading to misdiagnoses.”
POSTSCRIPT
In Maryland, House Bill 823 “died” in committee. It was titled “Generative Artificial Intelligence – Training Data Transparency.” The official synopsis was: “Requiring a developer of a generative artificial intelligence system, on or before January 1, 2026, and before the developer releases or substantially modifies a certain generative artificial intelligence system, to publish on the developer’s website documentation detailing the data used to train the generative artificial intelligence system.” For more information on legislative efforts, please see Bill to Create A.I. Evidence Clinic Pilot Program Was Vetoed in MD (Sep. 15, 2025).