Sony Music Publishing and Warner Chappell Music have taken Anthropic to federal court, opening another front in the increasingly messy fight over where AI companies get the enormous amounts of copyrighted material used to train their models.
The lawsuit, filed on August 28 in the US District Court for the Northern District of California, names Anthropic alongside co-founders Dario Amodei and Benjamin Mann. The publishers allege that copyrighted musical works were obtained without permission through large-scale torrenting, scraping and downloading, before being used in connection with the development of Anthropic’s Claude AI models.
At the centre of the case is an important distinction that could have wider consequences for generative AI. The argument isn’t simply about whether training an AI model on copyrighted material qualifies as fair use. Sony and Warner are attacking how Anthropic allegedly acquired some of that material in the first place.
The complaint alleges that Anthropic obtained millions of files from sources including Library Genesis and Pirate Library Mirror. According to the publishers, some books in those collections contained protected lyrics and sheet music, including compositions such as “Eye of the Tiger,” “September,” “Hallelujah,” “Ain’t No Mountain High Enough” and “All I Want for Christmas Is You.” The allegations have not been proven in court, and Anthropic has said it disagrees with the claims and intends to defend itself.
The potential financial exposure is significant. The publishers are seeking statutory damages that could reach $150,000 for each work found to have been willfully infringed. With the complaint involving tens of thousands of compositions, the theoretical damages could quickly reach billions of dollars, although any eventual award would depend on what the plaintiffs prove and how the court applies copyright law.
This isn’t Anthropic’s first encounter with the music business. Universal Music Publishing Group, Concord and ABKCO originally sued the company in 2023 over allegations involving song lyrics. A separate case filed in January 2026 covers more than 20,000 works, while other publishers have subsequently brought their own actions. Anthropic has argued in related litigation that using copyrighted works to train AI can constitute fair use.
That makes the latest Sony and Warner case particularly interesting. AI companies have increasingly built their legal defence around the transformative nature of model training, but obtaining copyrighted material through alleged piracy potentially creates a separate problem from what happens after that data enters a training pipeline.
The lawsuit therefore reaches beyond Claude or even music. Generative AI requires enormous datasets, and courts are gradually being asked to decide not only what machines may learn from, but how developers are allowed to obtain the material they learn from. If acquisition method becomes as legally important as model training itself, AI companies may find that the provenance of their datasets matters considerably more than it did during the industry’s rapid expansion.

