Court Documents Reveal: Microsoft Executive Called AI Data Scraping "the Greatest Theft of Labor in History"
Published: 17 September 2026
A newly unsealed set of court documents casts an uncomfortable light on how tech giants in the artificial intelligence sector have handled copyright issues behind closed doors. According to TechCrunch, internal communications between Microsoft representatives show that one company executive described OpenAI's scraping practices as "the greatest theft of labor in human history."
What the Unsealed Documents Show
According to the cited source, the court filings stem from a broader lawsuit concerning the unauthorized use of editorial content to train artificial intelligence models. The documents suggest that both Microsoft and OpenAI extracted content from behind The New York Times's paywall, subsequently using it to build datasets for training their algorithms.
TechCrunch further notes that some Microsoft employees warned internally that such practices could have serious consequences for the media industry, undermining the business models of publishers who rely on subscriptions and organic traffic to generate revenue.
A Paradox of the AI Industry
The situation highlighted by these documents raises a fundamental question for the entire industry: how can technology companies develop high-performing language models without compromising the rights of original content creators? The fact that a Microsoft executive used the word "theft" to describe the practices of a strategic partner — OpenAI, in which Microsoft has invested billions of dollars — reveals an internal tension between technological ambition and awareness of ethical and legal risks.
This revelation comes amid a wave of similar lawsuits filed by numerous publications and authors against AI companies, accusing them of using copyrighted material without authorization to train generative models.
What It Means for the Business World
For Romanian and international companies integrating AI solutions into their operations, this case serves as a warning about the need for greater due diligence when selecting technology providers. The provenance of the data used to train AI models is becoming an increasingly important criterion in assessing the legal and reputational risks associated with adopting artificial intelligence in business.
The litigation is still ongoing, and its outcome could significantly influence how the AI industry approaches copyright and data collection issues going forward.
Source
TechCrunch →844-ai.ro reports based on the source above. Editorially synthesized article, with attribution.
Comments
Loading discussion…
Checking your session…
Subscribe to our newsletter
Get the most important AI news once a week, straight to your inbox.