AI and Data Theft: What ‘Theft’ Means in the Age of Scraped Data

AI and Data Theft: What 'Theft' Means in the Age of Scraped Data

The email arrived minutes after GPT-6 Astra debuted: JoyIn’s CEO Guo Renjie called OpenAI out by name and promised a lawsuit. You could feel the tweetstorm forming—lawsuits, federal advisories, and a chorus of companies pointing fingers. I read the filings, the open letter, and the CISA advisory and realized this argument is not about code; it’s about ownership narratives.

JoyIn accused OpenAI publicly, and the accusation landed like a dropped match

JoyIn’s Guo Renjie published an open letter claiming OpenAI distilled Aether and copied his site’s “cosmic” design for GPT-6 Astra.

That complaint is notable for being the first time a Chinese AI developer has publicly accused an American firm of distillation. You should care because the charge flips a familiar script: Western firms have long accused Chinese rivals of illicit distillation; now a Chinese startup is pointing the same flashlight back.

What does distillation mean here? At its simplest, distillation uses a large model as a teacher to train a smaller student model. The technique is common, technically legal, and morally fuzzy—especially when models are trained on scraped web content without explicit consent.

Can companies sue over AI training data?

Yes—but winning is messy. Artists, news publishers, and even Apple have sued OpenAI for using copyrighted material and allegedly misappropriating trade secrets; OpenAI counters with claims of fair use. Courts are still sorting out whether scraping public websites to train models crosses the line from research to theft.

What is model distillation and why does it matter?

Distillation is a speed-and-efficiency play: you compress knowledge from a bulky teacher into a nimble student. The controversy arises when a company relies on a rival’s proprietary outputs as its teacher—and when the teacher itself was trained on copyrighted or private data.

The U.S. government’s advisory named companies and raised the stakes

CISA published an advisory saying firms such as DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.ai were running widescale extraction campaigns.

The agency described these efforts as industrial-scale knowledge distillation—arguing they were central to product development and not merely supplemental. Anthropic’s report added a twist: Moonshot may have been routing Claude’s outputs through its Kimi interface to mislead users about origin. These claims string national security, competition, and intellectual property into a single knot.

Artists, professors, and Apple have all filed suits; the complaints are piling up

Since ChatGPT launched, creators and publishers have sued over scraped content; Apple later accused OpenAI of stealing company secrets for hardware work; even an NYU professor worried his Codex prompts influenced later research results.

OpenAI frequently points to fair use as its defense, but you and I both know legal defenses are not the same as public trust. When the same industry cries foul about others’ distillation while depending on mass scraping, the moral claim loses clarity.

On the product side, everything is blurred: APIs, branding, and provenance

Companies are racing to ship models, polish brand identities, and lock in customers ahead of IPOs and major funding rounds.

I watch firms such as OpenAI and Anthropic prepare for what many expect will be record-setting public offerings while they also fight lawsuits and regulatory scrutiny. Startups and giants drift like pirates on a digital ocean, grabbing code and culture and calling it innovation.

Technical fixes exist—watermarking outputs, signed provenance, stricter API rate limits, and provenance metadata—but they require industry coordination and regulatory teeth. You should be suspicious of any single fix pitched as a cure-all; companies will trade control for speed whenever the market favors growth.

Society is asking a simple question; the legal system is struggling to answer it

Courts and regulators are now the referees in an argument they are ill-equipped to resolve quickly.

AI models are mirrors that reflect an internet’s worth of stolen light. If you build or buy AI, you will soon have to answer whether provenance matters more than performance, and who pays when provenance is murky.

So who gets to declare something stolen when the whole toolset was forged from content most of us never consented to—users, creators, governments, or the companies themselves?