Thursday, September 3, 2026

Sony and Warner Join Forces to Sue Anthropic: The Data You Feed Your AI Now Comes With a Price Tag

Anthropic has been sued by Sony, Warner, and book authors for massive unauthorized use of copyrighted music and books in training its AI models, facing billions in damages—a clear sign that data compliance has turned from a technical issue into a make‑or‑break business risk and cost for AI companies.

11 min read
Sony and Warner Join Forces to Sue Anthropic: The Data You Feed Your AI Now Comes With a Price Tag

Anthropic, which drew heated debate earlier over its AI book-burning plan, has once again found itself in the spotlight.

The reasons come down to two things. First, Anthropic is planning to go public and intends to file its IPO prospectus after Labor Day. Second, on the eve of that IPO, Sony and Warner jointly sued Anthropic for copyright infringement.

01

The complaint names two defendants in black and white: CEO Dario Amodei and co‑founder Benjamin Mann.

The plaintiffs allege that Anthropic obtained “tens of thousands” of musical works without authorization through BitTorrent, web scraping, and third‑party data collection.

As plaintiffs, Sony and Warner are seeking statutory damages of up to $150,000 per infringed work, plus up to $25,000 for each removal or alteration of copyright management information.

It is worth noting that Benjamin Mann once described pirate sites as “sketchy AF”—which, to translate, means this stuff is suspicious as hell. But the outcome is clear: he still participated in the relevant downloads.

The complaint explicitly states that the CEO personally ordered and approved the pirated downloads, and that Benjamin Mann personally carried them out.

Ironically, Anthropic has long positioned itself as a “safety‑first” company and defines itself as an “AI safety and research company.” Even Claude has carried the banner of being the “safety camp” in the large‑model industry. Yet today, both founders find themselves on the defendant’s bench, facing piracy charges.

Sony and Warner have gone further, calling the conduct one of the largest, most brazen, and most sustained thefts of intellectual property in Anthropic’s history.

And this is far from the first time Anthropic has found itself caught in an infringement storm.

02

In 2024, a group of novelists and nonfiction authors filed a class‑action lawsuit, alleging that Anthropic systematically downloaded and stored tens of thousands of copyrighted books from pirate libraries to train its Claude series models, without authorization or payment.

That case eventually settled for a record‑breaking amount of over $1.5 billion. The settlement fund is expected to cover around 480,000 to 500,000 eligible works, with a base payout of $3,000 per work. Additional terms include destroying the pirated datasets and certifying that they will not be used in any commercially deployed models.

On July 20, 2026—just a month ago—the court approved the $1.5 billion class‑action settlement between Anthropic and the authors and publishers.

Many will also remember that the same year, Anthropic launched the Panama Project, a secret operation costing tens of millions of dollars, which was exposed in January 2026. Through secondhand booksellers and other channels, Anthropic bought physical books published more than 22 years ago in bulk, used a hydraulic press to cut off the spines, scanned them page by page at high speed, and then fed the original books into a shredder—all to obtain high‑quality human text for training its AI models.

Anthropic documents once stated: “The Panama Project is our operation to destructively scan all the books in the world, and we don’t want the outside world to know about it.”

And even before Sony and Warner made their move, Anthropic was hardly innocent in music copyright—it had already accumulated four rounds of lawsuits. Universal Music Publishing Group (UMG), Concord Music Group, and ABKCO first sued in 2023, involving about 500 works. In January 2026, the same group of plaintiffs expanded the case to more than 20,000 works in a second lawsuit; in March they filed a third; and on August 17, Round Hill Music filed a fourth. None of these ended well for Anthropic.

Then, two days ago, Sony and Warner jointly filed a complaint, finally making Anthropic a co‑defendant against all three major music companies—Universal, Sony, and Warner—and putting it squarely at odds with the entire music industry.

This also raises some questions: Claude had already been fitted with guardrails for lyric output after multiple rounds of litigation, so why are rights holders still furious? Where exactly is the limit of fair use for training large models? And if online materials really do become priced training assets, who ultimately picks up the tab?

To trace these questions, we need to start with the birth of Anthropic and Claude.

03

In early 2021, a few former OpenAI employees founded Anthropic. By then, the AI large‑model race had already distilled an unspoken rule within the industry: to get good output from a model, you first have to feed it enough good text.

Randomly scraped web pages can teach a model to speak and respond, but books and lyrics, carefully crafted by creators, tend to provide much more stable structure and information. This is easy to understand: it’s like eating compressed biscuits every day versus having a private chef prepare your meals—over time, the difference shows in your body.

For a large model, books are certainly a well‑balanced meal, but if you want to feed someone else’s hard work to the model, you need to negotiate a price with the chef first.

Books can be bought and licensed, but the problem is obvious: it’s slow. The legal process for acquiring books includes, but is not limited to, contacting publishers, clearing rights, negotiating prices, and signing contracts. By the time you finish that process, you’re lagging far behind—roughly speaking, your model would barely be able to speak in complete sentences while competitors’ models could almost do backflips.

Copyright compliance is not impossible, but it costs time and money. When every competitor is racing for the training window, the temptation of a faster, cheaper data path is enormous.

Anthropic decided to take a shortcut—and unfortunately, that shortcut ended in court.

According to documents in the Bartz case, Anthropic began building a long‑term archive from its early days. In early 2021, co‑founder Benjamin Mann downloaded Books3, which contained 196,640 books. The court later found that it was entirely composed of unauthorized copies of copyrighted works, and determined that Mann knew the source was unauthorized. Two years later, Books3 was taken down by its host following a copyright complaint.

In June of the same year, Anthropic downloaded at least 5 million more books from LibGen, a site known as the “shadow library”—or, to put it plainly, a pirate database. In July 2022, Anthropic downloaded at least another 2 million books from PiLiMi, whose full name is Pirate Library Mirror—in other words, a mirror of a pirate library.

We already know the outcome: Anthropic paid $1.5 billion for that shortcut. But its digital trail had already been left behind, and this case has become strong evidence for Sony and Warner’s allegations that the founders participated in and approved company decisions.

04

Of course, being named in a complaint does not automatically establish personal liability. Sony and Warner still need to prove what the two did, what they knew, and how those actions connect to the infringement of specific works.

But for every AI entrepreneur, the signal from these lawsuits is clear. Training data is no longer just a spreadsheet or a pile of documents for the tech team; it is becoming an asset that boards and founders must sign off on.

Anthropic, for its part, seems to have made the same mistake that all AI large models tend to make, and it still has one possible defense: “fair use.”

Let’s define the term. Fair use is a legal doctrine that allows the use of copyrighted works without the rights holder’s permission in certain situations, such as commentary, research, search, and transformative creation. Large model companies have long argued that the model is learning patterns of language, not selling the original books to users, and that training is a highly transformative use. This argument is not without basis.

The Bartz case also brought Anthropic a half‑good news. The judge found that, based on the evidence in that case, copying specific books into a training set for training a large model was fair use; purchasing physical books, cutting them apart, scanning them into digital versions, and destroying the originals, as long as the digital versions were not distributed publicly, also fell within fair use.

The bad news came in the other half: the court held that copying from pirate databases is not protected by fair use, and the money and procedural hassle Anthropic saved are not legitimate justifications. In other words, training a model may be transformative, but that doesn’t mean you can steal the ingredients first. You can study how the chef cooks, but you can’t walk off with the chef’s raw materials.

05

The music publishers who are now the aggrieved parties are taking exactly this approach. They are not just focusing on the few lyrics Claude outputs to users; instead, they are breaking down the entire data chain into a series of fatal questions:

Where did your works originally come from? Have they been permanently retained? Were they included in the training input? What copies were made during training? Can the model reproduce the lyrics? Were copyright management information—such as work titles and authors—removed during data cleaning?

Even though Claude long ago stopped providing full lyrics, stopping what Claude says today is not the same as stopping what Claude consumed yesterday. Guardrails don’t explain whether training materials were licensed, whether rights information was stripped during cleaning, or whether reproducible content remains inside the model. And we all know the truth: those issues lie upstream of the guardrails.

So we can understand that the rights holders’ anger is not limited to whether AI can spit out lyrics, but whether AI companies can prove the supply chain for all their data.

But the business spawned by AI won’t stop. Even as lawsuits multiply, the legal entry point has officially opened as a business. At the end of 2025, music‑tech company KLAY announced it had reached AI licensing agreements with the record labels and publishing divisions of Universal, Sony, and Warner. KLAY said its large music model is trained entirely on licensed music, and that participating artists and songwriters will be recognized and compensated. The three major labels also emphasized that Klay not only respects artists, songwriters, and rights holders, but also provides a complete mechanism for AI licensing, usage, and revenue distribution.

Soon after, music data platform Musicmatch followed Klay’s lead and announced AI licensing deals with the three majors, with access to a catalog of over 15 million musical works. While AI large‑model companies are still entangled in lawsuits, the door to the licensing market has already swung open.

06

It looks perfectly legal and above board, but at the same time, a bill is slowly taking shape.

First, there are the licensing fees, which are fairly uncontroversial: whether for books, lyrics, or even images and news, once these data sources have tradable licenses, data that was once invisible on the books will gradually become explicit.

Second, data governance costs: model providers need to understand the origin and path of each batch of data. How long this kind of training will take, and how many versions it will require, remains unknown.

Third, model engineering: when previously used, disputed data without clear rights is deemed unusable, the consequences go far beyond hitting delete. The company must figure out which data mixtures these sources entered, which model versions they affected, and then decide whether to fine‑tune or retrain. The compute and time costs involved are enormous and hard to estimate.

Fourth, litigation preparedness and customer indemnification: after all, no one wants to pay for a model that could pass a lawsuit on to them at any moment. So intellectual property indemnification has started to move from a contract appendix to a product selling point. Anthropic’s current commercial terms are quite direct: if a third party alleges that a customer’s paid, compliant use of the service infringes intellectual property, Anthropic will defend the customer and cover any damages from a court judgment or approved settlement. In other words, if an enterprise uses Claude, Anthropic will bear some of the copyright risk on your behalf.

Fifth, and also a major cost, is market access: once the licensing market opens, large enterprises will quickly abandon smaller models whose data records and indemnification terms are unclear. The big players have the money to buy licenses and the capital to bear liability. Smaller model companies, meanwhile, have to chase performance while covering legal costs, so their room to survive will shrink dramatically. In the end, the only ones likely to smile are the few companies rich enough to buy data rights.

07

Anthropic’s latest dust‑up is just one copyright lawsuit, but for many enterprise buyers and users, it can also be seen as a renegotiation of how AI production costs are allocated.

From this, we can foresee that when the AI model war enters its next phase, large models will still tout compute, speed, and price as core selling points. But the difference is that when they enter a company’s core business, they will also have to compete on who can produce a supply chain record that can withstand scrutiny from both courts and customers.

Capability may determine the ceiling of an AI product, but only a written record of data provenance and a contract guarantee in black and white are the best chips for whether it can survive within an enterprise.

Disclaimer: The information provided in this article is for general informational purposes only. While we strive for accuracy, NewsHub makes no representations or warranties about the completeness or reliability of the content. Always verify important information from multiple sources.

Category:Finance
Share:

hanno

NewsHub editorial team member. Dedicated to providing you with high-quality, fact-checked news coverage.

More in Finance

Overtaking SpaceX? The Biggest IPO in History on the Horizon?

Overtaking SpaceX? The Biggest IPO in History on the Horizon?

Finance

Overtaking SpaceX? The Biggest IPO in History on the Horizon?

Anthropic plans to acquire AI optimization firm Decart for $6 billion and may go public at a record $2 trillion valuation, highlighting that in the AI era, tech giants are pricing scarce technology and R&D time ahead through premium acquisitions.

Gujie13 min
Giving Robots "Spatial Intuition" – Ommo Technologies Secures Tens of Millions of USD in Series A Funding

Giving Robots "Spatial Intuition" – Ommo Technologies Secures Tens of Millions of USD in Series A Funding

Finance

Giving Robots "Spatial Intuition" – Ommo Technologies Secures Tens of Millions of USD in Series A Funding

Ommo Technologies, with its proprietary permanent‑magnet positioning technology, delivers high‑precision, occlusion‑resistant, and scalable “spatial intuition” for robots and physical AI, and has just closed a tens‑of‑millions‑of‑USD Series A round to accelerate deployments in healthcare, embodied intelligence, and beyond.

yong9 min
Porsche Launches Panamera "Era" Edition in China with Starting Price of 998,000 RMB

Porsche Launches Panamera "Era" Edition in China with Starting Price of 998,000 RMB

Finance

Porsche Launches Panamera "Era" Edition in China with Starting Price of 998,000 RMB

Porsche has launched the Panamera Era Edition in China with a starting price of 998,000 RMB, aiming to boost sales through price cuts and more standard equipment, but this move has triggered knock‑on effects such as brand dilution and a sharp drop in used‑car values, reflecting Porsche’s dilemma between protecting sales volume and protecting profit margins in the current market environment.

anna13 min