Newly unsealed court documents have revealed significant internal concerns at Microsoft and OpenAI regarding their use of copyrighted news content to train artificial intelligence systems. According to TechCrunch, a Microsoft executive privately characterized OpenAI’s data practices as “the largest theft of labor in human history.”
The filings show that both companies scraped paywalled content from news outlets, including The New York Times, and built datasets from this material, according to TechCrunch. The NYT Technology reports that the documents revealed concern within both organizations over the use of millions of news articles to develop their AI systems. According to TechCrunch, internal warnings indicated these practices would “gut publishers.”
The revelations provide insight into the tension between AI companies’ data needs and content creators’ rights, with executives at these companies apparently aware of potential legal and ethical issues even as they proceeded with the practices. The documents emerged as part of ongoing litigation between technology companies and content publishers over AI training data.