Skip to content

AI Training Copyright Risk: What New Filings Reveal

Juwel Rana

By Juwel Rana · CEO & Founder

1,078 views
A lively railway scene in Bangladesh with trains, people, and lush greenery.

Newly unsealed court filings in The New York Times' copyright lawsuit against OpenAI and Microsoft show a Microsoft director describing the industry's own data practices in blunt terms.

He called it "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history."

That admission, made in a January 2023 internal memo and only unsealed this month, points to a real AI training copyright risk for any company training or buying AI built on content it doesn't own.

The memo's author, Brent Hecht, is Microsoft's director of applied science. He also argued that letting the companies' fair use defense succeed would "make a complete mockery of the idea of 'fair use.'"

The Times now cites that line as proof OpenAI and Microsoft knew exactly what they were doing.

What the Unsealed Filings Actually Show

The Times sued OpenAI and Microsoft in 2023 over training ChatGPT and Copilot on its journalism without permission or payment.

A judge isn't expected to rule on whether it goes to trial until 2027, based on the newly unsealed court motion.

What that motion reveals goes beyond Hecht's quote. Microsoft's own data, cited in unsealed court documents reported by TheWrap, shows click-through rates from Bing Chat to the Times' site fell as much as 93% below what traditional Bing search sent.

According to TechCrunch's review of the filings, one OpenAI dataset alone held more than 91,000 copies of articles from the Times, the Daily News and the Center for Investigative Reporting.

An OpenAI employee described finding "a hack to get around nytimes paywall." Cofounder Greg Brockman's reply was two words: "ah nice."

The AI Training Copyright Risk Every Business Now Faces

None of this is abstract for companies outside the AI industry. Microsoft CEO Satya Nadella testified that paywalled material "should be licensed by anyone who wants to use it…for grounding or training."

He added that he'd have required OpenAI to retrain its models had he known paywalled Times content was in the training data. That's a founder-level admission that scraping without a license is a decision executives already regret.

Any company feeding its own model on scraped or paywalled material, or buying a vendor's model without asking where the training data came from, carries the same exposure now undercutting OpenAI and Microsoft's fair use defense in court.

Executives at the companies being sued called their own practices theft in private. A smaller company copying the same approach has little cover to claim it didn't know better.

Protecting Your Company Before You Train or Buy AI

Commuter Train 811 captured at Saga Station on a rainy day in Japan.

Photo by Glen Zi 加侖子 on Pexels

The practical response isn't complicated. Ask any AI vendor exactly what data trained the model, and get the answer in writing. Read the terms of service on any site your own tools scrape or query, especially if you're building automations with AI and automation that touch third-party content.

If your business depends on organic search traffic the way publishers do, this case is worth watching too. Chatbots that answer directly instead of sending readers to the source create the same substitution problem for any company whose digital marketing and growth strategy leans on organic traffic.

We've covered how to manage that exposure in how to block AI training without blocking Google.

For companies building or fine-tuning their own models, licensing data outright beats betting on a fair use defense that looks shakier by the month. Our breakdown of what the OpenAI copyright case means for AI training explains what that looks like in practice.

The Trump administration filed a brief in September 2026 backing OpenAI's fair use position, so the legal outcome is far from settled.

But the unsealed testimony already answers the question that matters most for your own AI decisions: the people running these companies knew the risk before regulators or courts did, and built anyway.

Do your own diligence before you follow the same path.

Cover photo by Nirjon Nakib on Pexels

Latest Blog

A smartphone with a shopping cart depicting the concept of online shopping in a colorful studio setup.Ecommerce • Digital Strategy

Building an Ecommerce Online Presence From Scratch

A new store isn't competing for one moment on a website anymore. Here's how the search, social and review data from 2026 should shape where you spend the first few months.

Read More
Focused view of programming code displayed on a laptop, ideal for tech and coding themes.Travel & Hospitality • SEO

Travel and Hospitality Schema Markup Explained

Most hotel and travel sites either skip structured data or build it wrong. Here's what schema.org actually requires, and where Google's own rules just changed.

Read More
Young man wearing eyeglasses working on laptop in vibrant digital agency office.E-commerce • Digital Marketing

Freelancer vs Agency Ecommerce: Or Hire In-House?

Freelancers, agencies and in-house hires solve different problems for an online store. Here's what an in-house hire really costs and when it's worth switching models.

Read More
Macro photography of color palette code in a programming environment.Manufacturing • No-Code

No-Code vs. Custom Development for Manufacturing Startups

No-code platforms are the fast, cheap first move for a manufacturing startup, until an ERP built in 2006 or an AS9100 audit trail enters the picture. Here's where the no-code vs. custom development line actually falls.

Read More
A minimalist image featuring the words 'Branding' and 'Marketing' on a white background, ideal for digital marketing themes.Healthcare • Branding

How Healthcare Brand Identity Builds Patient Trust

Patients decide how much to trust a practice before they read the About page. Here's what actually shapes that first impression, and what two real health system rebrands changed to earn it back.

Read More
Futuristic abstract digital render depicting geometric shapes in vibrant colors.SEO • AI

How to Block AI Training Without Blocking Google

Cloudflare's new default rules bundle Googlebot with AI training crawlers unless you change a setting. Here's what actually changed on September 15 and how to keep search access while opting out of training.

Read More

Subscribe to our newsletter

Offers, insights and updates — a couple of times a month, never more.