New York Times lawsuit against OpenAI: Where is the case headed?
The widely followed legal battle between The New York Times, OpenAI and Microsoft over whether news content can legally be used to train artificial intelligence models has entered a critical stage. After nearly three years of litigation, the case has recently taken several dramatic turns. On one hand, previously confidential court documents have come to light; on the other, the US government has formally sided with the two technology companies.
The New York Times filed the lawsuit in federal court in Manhattan in December 2023. The newspaper alleged that OpenAI and Microsoft had used its news reports, investigative journalism and opinion pieces without permission or payment to develop competing commercial products. In March 2025, Judge Sidney Stein dismissed some of the claims but allowed the core of the lawsuit to proceed. Lawsuits brought by several other media organisations and writers, including the Daily News, the Center for Investigative Reporting and The Intercept, were later consolidated with the case.
What the confidential documents revealed
On September 17, an unredacted version of a document submitted by The New York Times was made public. It contained internal comments by executives at the two companies.
The document quoted a 2023 internal memo by Brent Hecht, Microsoft’s director of applied science, describing the large-scale collection of content as “an astounding theft of unprecedented scale” and “the largest labor theft in human history.”
However, the context of the full statement was somewhat different. The memo was actually predicting that millions of people around the world would soon view the collection of content by large AI models in those terms. In other words, the statement was a forecast of public sentiment. This has become a point of dispute between the two sides over how the quotation should be interpreted.
The plaintiffs’ filings also alleged that large-scale content collection was carried out by bypassing paywalls and other restrictions on paid content, and that copyright notices were removed from training datasets.
According to the documents, OpenAI’s mid-training dataset alone contains more than 91,692 copies of content published by The New York Times, the Daily News and the Center for Investigative Reporting. Another dataset created from Common Crawl contains more than two million documents.
The plaintiffs argue that the two companies themselves understood that AI products could create an existential crisis for news organisations and trigger a “doom loop” for the web.
It should be noted, however, that much of the information made public came from the plaintiffs’ own legal filings. The underlying source documents on which those arguments are based remain sealed under court orders. Some portions of the filings, particularly information involving financial figures, also remain redacted.
The defendants’ arguments
OpenAI and Microsoft have denied the allegations from the outset.
In a court filing, Microsoft argued that existing copyright law does not give rights holders the authority to prevent transformative technologies. OpenAI has maintained that its technology makes the world’s information accessible to anyone and that copyright law should not stand in the way of such technological progress.
On September 4, all three parties formally submitted arguments seeking summary judgment. The New York Times has asked the court to determine liability. If the judge does not rule in its favour, the case could proceed to a jury trial.
The government’s position
A major turning point came on September 2, when the US Department of Justice filed a statement with the court supporting the positions of OpenAI and Microsoft.
The Justice Department argued that training AI models on copyrighted written material cannot, in general, be considered copyright infringement. It also warned that ruling in favour of the plaintiffs could pose a threat to national security and give foreign competitors an advantage.
This is the US government’s first formal position on the copyright implications of using copyrighted material to train AI models.
A spokesperson for The New York Times responded by saying the government had chosen to side with a handful of trillion-dollar AI companies at the expense of American creators.
Why the case matters
Legal analysts say most copyright lawsuits involving AI are still at relatively early stages and none has yet gone through a full trial. As a result, the outcome of this case could establish an important precedent for future litigation.
If the plaintiffs prevail, AI companies could be required to pay for the content they use to train their models, potentially putting significant pressure on their existing business models. If the defendants prevail, news organisations could find themselves with substantially less control over how their content is used and fewer opportunities to generate revenue from it.
The issue is also relevant to Bangladesh’s news industry. Content published by Bangladeshi news websites is also being used to train various AI models. Yet the country’s copyright law contains no specific provision addressing the use of content for AI training.
There has also been little substantial discussion within Bangladesh’s news industry about consent, compensation or licensing arrangements for such use. The outcome of the New York Times case could therefore become an important reference point as Bangladesh begins to confront similar questions over the use of journalistic content in the age of AI.
Leave A Comment