ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI compliance
AI Training-Data Copyright Red Line Tightens: OpenAI Allegedly Knew Pirated-Book Training Was Illegal

AI Training-Data Copyright Red Line Tightens: OpenAI Allegedly Knew Pirated-Book Training Was Illegal

AI compliance • Admin • • 2 views

The copyright red line for AI training data is tightening. On September 21, 2026, the Authors Guild disclosed internal documents cited in the plaintiffs' latest brief in Alter v. OpenAI and Microsoft: executives at OpenAI and Microsoft knew years ago that training models on books from the pirate library LibGen carried legal risk — they did it anyway, and in the summer of 2022 tried to erase the traces.

The class action is being heard in the U.S. District Court for the Southern District of New York (Case No. 1:23-cv-08292), part of multidistrict litigation against the two companies. The plaintiffs include the Authors Guild and more than a dozen authors, among them George R.R. Martin and John Grisham. On September 17, 2026, the plaintiffs filed a motion for partial summary judgment (Docket 1982) along with a 162-page statement of facts (Docket 1987), which the Guild then released to the public.

What happened: from downloading pirated books to "Project Clear"

According to internal documents cited in the plaintiffs' brief, the timeline looks like this:

In 2018, OpenAI downloaded roughly 117,500 books from LibGen, then torrented about 35TB of data — which the brief says amounted to LibGen's entire collection of 4.6 million books at the time. In April 2019, when Sam Altman showed an early version of GPT-3 to Bill Gates, the accompanying document stated the company had "added another ~11B words from Library Genesis (LibGen)"; two months later, Microsoft invested $1 billion in OpenAI. In the GPT-3 paper, the datasets were called "Books1" and "Books2."

More critical is the evidence of knowledge. In May 2020, OpenAI policy director Jack Clark wrote internally that the models "will make people unemployed" and that "there will be a point where a bunch of artists express worry about what we're doing here and we'll likely ignore their concerns and release anyway." Tarun Gogineni, who joined OpenAI in 2022 to improve the models' writing quality, dismissed authors' complaints that the "datasets are stolen" as an "acceptable economic disruption," and mused about having GPT finish the last two books of A Song of Ice and Fire.

Microsoft is also swept into the knowledge allegations: the brief says that during the April 2019 demo, Altman and Dario Amodei — then OpenAI's research director — disclosed the LibGen use to Gates and others. Amodei described LibGen as "a bit sketchier" as a training set; researcher Sam McCandlish was worried about "optics" — that "'openai uses copyrighted data from sketchy russian website' showing up on Hacker News would be unfortunate."

In the summer of 2022, OpenAI launched a cleanup operation code-named "Project Clear" and deleted its LibGen files. Slack records from June 15, 2022 show someone asking what to do about "libgen" mentions being "all over google docs/slack/github"; that evening, VP of Research Bob McGrew replied: "Given how much OpenAI is in the news, now is the right time to excise Libgen from our systems and storage." Earlier, Judge Ona Wang had already rejected OpenAI's privilege claims over these Slack records and ordered the company to turn over chat logs from the "project-clear" and "excise-libgen" channels.

Who is affected: AI developers first, authors and publishers gain leverage

The most direct impact falls on AI model developers and training-data owners. Past debates over training-data copyright mostly centered on whether fair use covers training; this brief moves the battleground to whether downloading and storing pirated books is itself infringement — and whether the companies acted with knowledge. If knowledge is established, the room for a fair-use defense shrinks dramatically. In a separate AI copyright case, the U.S. Ninth Circuit dealt with whether AI-generated outputs remove copyright management information; this case is about the legality of training-data sources — a different battleground.

Authors and publishers are on the other side. The Authors Guild has 18,000 members, and the plaintiffs include several bestselling authors. Notably, just days before these filings were disclosed, similar documents were unsealed in The New York Times v. Microsoft and OpenAI, showing that training on millions of news articles poses an "existential threat" to the newspaper business — the two fronts are converging.

But boundaries must be drawn: everything above comes from the plaintiffs' brief and the internal documents it cites — one side's claims. The court has not ruled on whether infringement occurred or whether fair use applies. The brief seeks partial summary judgment; both sides will file more briefing in the coming months, with a hearing expected in early 2027.

What to do: audit data lineage first, don't delete records first

For teams that have trained or are training models, this case points to four actionable steps:

First, check whether training datasets contain pirated sources like LibGen or Z-Library. Run a full data-lineage audit documenting each batch's source, acquisition method, and licensing status.

Second, negotiate licenses where possible, and soon. Anthropic already paid $1.5 billion to settle a similar pirated-book training suit (Bartz v. Anthropic) — "build first, license later" may cost far more than licensing upfront.

Third, remove discovered pirated data with a paper trail — but don't copy "Project Clear." If deletion is read as covering up evidence, it becomes proof of knowledge instead. Here, the judge ordered the turnover of all related Slack records for exactly that reason.

Fourth, mind internal communications. Discussions of data sources in Slack, docs, and paper drafts are all discoverable. The real fix is compliance before the fact, not optics management after it.

Where the risks are: three numbers and one date

The first number is $150,000 — the statutory damages cap per work for "willful" infringement under U.S. copyright law. The plaintiffs' brief is built around knowledge precisely to reach that tier: 117,500 books times $150,000 is an astronomical figure.

The second number is $1.5 billion — the Anthropic settlement. In a parallel case, a federal judge confirmed that illegally downloading pirated books is standalone infringement, even if the company later bought legitimate copies. If that logic carries into this case, OpenAI's fair-use defense gets much harder.

The third risk is timing: the hearing is expected in early 2027, with more briefing due in the coming months. For companies using third-party model APIs, the risk is indirect but real — if the training data is found unlawful, the supplier's compliance defect can travel down the supply chain.

One final note on scope: this lawsuit is fought under U.S. copyright law, and LibGen has been on the U.S. Trade Representative's "Notorious Markets" list since 2017. Training-data rules differ across jurisdictions — the trajectory of a U.S. court should not be mapped directly onto global operations.

For AI companies, this case upgrades training-data compliance from "legal advice" to "discovery-grade" risk: where the data came from, who knew, and when they knew it can all become exhibits.

Recommended Tools

More