POLICY · REGULATION · #573
Unredacted NYT filings show Microsoft/OpenAI execs called AI scraping 'theft' and warned models threaten publishers
Unredacted material from The New York Times' copyright lawsuit against OpenAI and Microsoft quotes internal statements saying AI training practices amounted to ‘‘theft’’ and that models are ‘‘substitutive’’ or an ‘‘existential threat’’ to publishers. The filings allege large-scale scraping (including millions of NYT URLs in Common Crawl-derived data), internal projects sharing training datasets (e.g., Project Mango/Taxi and delivery of GPT-3 training data to Microsoft), and Microsoft metrics showing Copilot reduced NYT click-throughs by as much as 93%.
KEY POINTS
- Unredacted material from The New York Times' copyright lawsuit against OpenAI and Microsoft quotes internal statements saying AI training practices amounted to ‘‘theft’’ and that models are ‘‘substitutive’’ or an ‘‘existential threat’’ to publishers.
- The filings allege large-scale scraping (including millions of NYT URLs in Common Crawl-derived data), internal projects sharing training datasets (e.g., Project Mango/Taxi and delivery of GPT-3 training data to Microsoft), and Microsoft metrics showing Copilot reduced NYT click-throughs by as much as 93%.
- These admissions and dataset details could undercut fair-use defenses, influence the outcome of major copyright litigation, and push changes to licensing, product design, or regulation for AI training data.
WHY IT MATTERS
These admissions and dataset details could undercut fair-use defenses, influence the outcome of major copyright litigation, and push changes to licensing, product design, or regulation for AI training data.