Tech Meridian ← LIVE FEED
RU

POLICY · REGULATION · #573

Unredacted NYT filings show Microsoft/OpenAI execs called AI scraping 'theft' and warned models threaten publishers

Unredacted material from The New York Times' copyright lawsuit against OpenAI and Microsoft quotes internal statements saying AI training practices amounted to ‘‘theft’’ and that models are ‘‘substitutive’’ or an ‘‘existential threat’’ to publishers. The filings allege large-scale scraping (including millions of NYT URLs in Common Crawl-derived data), internal projects sharing training datasets (e.g., Project Mango/Taxi and delivery of GPT-3 training data to Microsoft), and Microsoft metrics showing Copilot reduced NYT click-throughs by as much as 93%.

KEY POINTS

  1. Unredacted material from The New York Times' copyright lawsuit against OpenAI and Microsoft quotes internal statements saying AI training practices amounted to ‘‘theft’’ and that models are ‘‘substitutive’’ or an ‘‘existential threat’’ to publishers.
  2. The filings allege large-scale scraping (including millions of NYT URLs in Common Crawl-derived data), internal projects sharing training datasets (e.g., Project Mango/Taxi and delivery of GPT-3 training data to Microsoft), and Microsoft metrics showing Copilot reduced NYT click-throughs by as much as 93%.
  3. These admissions and dataset details could undercut fair-use defenses, influence the outcome of major copyright litigation, and push changes to licensing, product design, or regulation for AI training data.

WHY IT MATTERS

These admissions and dataset details could undercut fair-use defenses, influence the outcome of major copyright litigation, and push changes to licensing, product design, or regulation for AI training data.

SOURCES & TIMELINE

1