Newly unsealed court filings in The New York Times' copyright lawsuit against OpenAI and Microsoft show a Microsoft director describing the industry's own data practices in blunt terms.
He called it "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history."
That admission, made in a January 2023 internal memo and only unsealed this month, points to a real AI training copyright risk for any company training or buying AI built on content it doesn't own.
The memo's author, Brent Hecht, is Microsoft's director of applied science. He also argued that letting the companies' fair use defense succeed would "make a complete mockery of the idea of 'fair use.'"
The Times now cites that line as proof OpenAI and Microsoft knew exactly what they were doing.
What the Unsealed Filings Actually Show
The Times sued OpenAI and Microsoft in 2023 over training ChatGPT and Copilot on its journalism without permission or payment.
A judge isn't expected to rule on whether it goes to trial until 2027, based on the newly unsealed court motion.
What that motion reveals goes beyond Hecht's quote. Microsoft's own data, cited in unsealed court documents reported by TheWrap, shows click-through rates from Bing Chat to the Times' site fell as much as 93% below what traditional Bing search sent.
According to TechCrunch's review of the filings, one OpenAI dataset alone held more than 91,000 copies of articles from the Times, the Daily News and the Center for Investigative Reporting.
An OpenAI employee described finding "a hack to get around nytimes paywall." Cofounder Greg Brockman's reply was two words: "ah nice."
The AI Training Copyright Risk Every Business Now Faces
None of this is abstract for companies outside the AI industry. Microsoft CEO Satya Nadella testified that paywalled material "should be licensed by anyone who wants to use it…for grounding or training."
He added that he'd have required OpenAI to retrain its models had he known paywalled Times content was in the training data. That's a founder-level admission that scraping without a license is a decision executives already regret.
Any company feeding its own model on scraped or paywalled material, or buying a vendor's model without asking where the training data came from, carries the same exposure now undercutting OpenAI and Microsoft's fair use defense in court.
Executives at the companies being sued called their own practices theft in private. A smaller company copying the same approach has little cover to claim it didn't know better.
Protecting Your Company Before You Train or Buy AI

Photo by Glen Zi 加侖子 on Pexels
The practical response isn't complicated. Ask any AI vendor exactly what data trained the model, and get the answer in writing. Read the terms of service on any site your own tools scrape or query, especially if you're building automations with AI and automation that touch third-party content.
If your business depends on organic search traffic the way publishers do, this case is worth watching too. Chatbots that answer directly instead of sending readers to the source create the same substitution problem for any company whose digital marketing and growth strategy leans on organic traffic.
We've covered how to manage that exposure in how to block AI training without blocking Google.
For companies building or fine-tuning their own models, licensing data outright beats betting on a fair use defense that looks shakier by the month. Our breakdown of what the OpenAI copyright case means for AI training explains what that looks like in practice.
The Trump administration filed a brief in September 2026 backing OpenAI's fair use position, so the legal outcome is far from settled.
But the unsealed testimony already answers the question that matters most for your own AI decisions: the people running these companies knew the risk before regulators or courts did, and built anyway.
Do your own diligence before you follow the same path.
Cover photo by Nirjon Nakib on Pexels





























