5 min read
The legal battle over how artificial intelligence companies build their models has taken another major turn. A coalition representing nearly 400 newspapers has filed a new lawsuit against Microsoft and OpenAI, accusing the two companies of using copyrighted news content without permission to train their AI systems.
The case adds to a growing list of copyright disputes facing the tech giants as publishers, authors, artists, and media companies push back against the way generative AI models are developed. At the center of the argument is a simple question with enormous implications: Can AI companies use online content to train their systems without compensating the people who created it?
According to the complaint filed in the U.S. District Court for the Southern District of New York, the publishers allege that Microsoft and OpenAI systematically copied their articles and other original works to develop AI chatbots.
The complaint claims the companies crawled newspaper websites, including content that sat behind paywalls and other restrictions, before storing that material on their own servers. The publishers argue that this was done without authorization and without providing compensation to the organizations that produced the content.

The coalition says the alleged actions have allowed the companies to generate billions of dollars in value while the news organizations that created the material received nothing in return.
As a result, the publishers are seeking statutory damages and injunctive relief, citing both copyright infringement and alleged violations of the Digital Millennium Copyright Act.
The lawsuit is unlikely to be the last chapter in the ongoing fight over AI training data. Microsoft and OpenAI have repeatedly argued in previous cases that copyright law does not explicitly ban the use of publicly available online material to train artificial intelligence models.
OpenAI has also maintained that its systems are built using publicly available information and that the company’s practices fall under the principle of fair use.
An OpenAI spokesperson recently reiterated that position, saying the company’s models empower innovation and are trained on publicly available data.
The issue has become one of the biggest unresolved legal questions facing the rapidly growing AI industry, with courts now being asked to determine whether existing copyright rules apply to this new technology.
Little-known fact: The first modern copyright law is widely considered to be the Statute of Anne, enacted in Britain in 1710 to give authors legal rights over their work.
The newspaper coalition says the dispute is about more than compensation. Publishers argue that the future of local journalism itself could be at stake if AI companies continue to use news content without entering licensing agreements.
Many local newspapers are already dealing with declining advertising revenue, shrinking staffs, and growing competition for readers’ attention. The rise of AI-powered search tools and chatbots has created new concerns that users may increasingly consume information through AI-generated summaries rather than visiting original news websites.
Some industry observers fear that such a shift could further weaken local news organizations that rely on traffic and subscriptions to survive.
Former New Jersey Attorney General Matthew Platkin, who is involved in the case, suggested that any eventual resolution should not benefit only the largest media companies while leaving smaller local outlets behind.
The legal fight comes at a time when the AI industry is facing another challenge: finding enough high-quality data to train increasingly sophisticated models.
Reports published in recent years have suggested that several leading AI companies, including OpenAI, Google, and Anthropic, have encountered limitations because high-quality training material is becoming harder to obtain.
That reality has intensified debates over licensing agreements and the value of professionally produced journalism. News organizations argue that their reporting represents a significant investment of time, expertise, and money, making it unfair for AI companies to use that work without compensation.
At the same time, AI developers warn that overly restrictive rules could slow innovation and make it more difficult to create useful new technologies.
The lawsuit from nearly 400 newspapers may ultimately become one of the most important copyright battles of the AI era. The outcome could influence how artificial intelligence companies gather data, how publishers protect their work, and whether new licensing models emerge between the technology and media industries.

If courts determine that AI companies must pay for the content used to train their systems, the decision could reshape the economics of artificial intelligence development. On the other hand, a ruling that broadly supports fair use could make it harder for publishers to control how their content is used in the future.
For now, the courtroom has become the latest battleground in the growing struggle between the companies building AI and the organizations creating the information those systems rely upon.
This article was made with AI assistance and human editing.
Don’t forget to follow us for more exclusive content.
If you liked this, you might also like:
This content is exclusive for our subscribers.
Get instant FREE access to ALL of our articles.
We appreciate you taking the time to share your feedback about this page with us.
Whether it's praise for something good, or ideas to improve something that
isn't quite right, we're excited to hear from you.
Lucky you! This thread is empty,
which means you've got dibs on the first comment.
Go for it!