Lawsuit filings from The New York Times and others call it 'unprecedented theft'; OpenAI and Microsoft say use was transformative and legally protected
Court documents made public this week contain allegations that OpenAI, the maker of ChatGPT, massively infringed on news organizations' copyrights by using more than 10 million articles without permission to develop its AI systems.
According to AFP, the documents — released Thursday (local time) — include claims from insiders that OpenAI's large-scale use of articles from The New York Times and other outlets to train ChatGPT amounted to theft.
Brent Hecht, head of applied science at Microsoft, which invested in OpenAI in 2019, said in the documents that OpenAI's unauthorized use of more than 10 million news articles was "an amazing theft at an unprecedented scale, the greatest labor exploitation in human history."
He also said OpenAI may have engaged in "inadvertent concealment" while reviewing the news articles it had collected.
According to the court documents, roughly one-third of the more than 10 million articles OpenAI used without authorization came from The New York Times.
The New York Times filed a copyright infringement lawsuit against OpenAI and Microsoft in federal district court in New York three years ago. The case has since been joined by Ziff Davis — the US media company that owns technology outlet CNET — the parent company of the American magazine Mother Jones, investigative news site The Intercept, and several local news organizations across the United States.
OpenAI began offering ChatGPT in late 2022 and has since signed content licensing agreements with various news organizations around the world. Concerns have persisted, however, that generative AI software would erode traffic to news websites.
AI companies attempted to address those concerns by adding citations with links to answers, but voices within the industry itself argued the measure was insufficient.
The court documents show that OpenAI's own engineers acknowledged the approach was largely cosmetic, saying "no matter how prominently we show links, users won't click on them."
OpenAI and Microsoft have pushed back against the allegations, arguing that their use of news content constitutes transformative use — adding value, purpose and function to the original works — and falls within the fair use doctrine recognized under copyright law.
The US Department of Justice filed a brief with the court earlier this month siding with OpenAI and Microsoft, citing scientific advancement, economic growth and national security.
OpenAI and Microsoft have asked the court for summary judgment — an immediate ruling in their favor — which, if granted by the presiding judge, would end the case without a full trial.
yckim6452@heraldcorp.com
