Incidents · New York Times Sues OpenAI for Copyright Infringement
AI INCIDENT

New York Times Sues OpenAI for Copyright Infringement critical

Date
December 27, 2023
Company
OpenAI
Product
ChatGPT
Category
Legal
Severity
CRITICAL

What Happened

On December 27, 2023, The New York Times filed a lawsuit in the Southern District of New York against OpenAI and Microsoft, alleging that the companies used millions of Times articles to train their AI models without permission or compensation. The suit claimed that ChatGPT and other models could reproduce Times content nearly verbatim, effectively creating a substitute for the original journalism.

The complaint included striking examples of ChatGPT outputting near-exact copies of Times articles when prompted, demonstrating that the training data had been memorized rather than merely learned from in a general sense.

Timeline

The NYT and OpenAI had been in licensing negotiations through much of 2023, but talks broke down before a deal could be reached. The Times filed suit on December 27, 2023, seeking billions in statutory and actual damages. The case moved through preliminary proceedings in 2024, with the court making several rulings on discovery and the scope of claims. As of early 2025, the litigation was ongoing with significant implications for both parties.

Impact

The NYT lawsuit became the most significant legal challenge to AI training data practices. Unlike individual creator lawsuits, the Times brought the resources, standing, and public profile to potentially force industry-wide changes. The suit raised fundamental questions about whether training AI models on copyrighted content constitutes fair use or requires licensing.

The financial stakes were enormous. The Times sought damages that could theoretically reach billions of dollars if statutory damages were applied per article. The case also had implications for every major AI company that trained models on web-scraped data.

Response

OpenAI argued that training on publicly available content was fair use and that the examples of verbatim reproduction were artificially induced through specific prompting techniques. The company maintained that AI models learn patterns from training data rather than storing and reproducing copyrighted content.

Microsoft largely aligned with OpenAI’s position while emphasizing its own fair use arguments. Both companies continued to pursue licensing deals with other publishers while litigating the NYT case.

Lessons Learned

The NYT lawsuit crystallized the unresolved legal questions at the heart of the AI training data debate. It demonstrated that the industry’s implicit assumption — that web scraping for AI training was legally permissible — was far from settled law. The case forced every major AI company to reassess its training data practices and pursue proactive licensing deals.

The incident also revealed the strategic importance of content ownership in the AI era. Publishers who had seen their content devalued by the internet economy suddenly found themselves holding leverage in AI negotiations. The NYT’s willingness to litigate rather than accept unfavorable licensing terms showed that premium content creators were prepared to fight for compensation.