Hacker News new | ask | show | jobs
by tristanj 10 days ago
The main story is what isn't being talked about. Chinese labs exfiltrated trillions of tokens of high-quality output from Anthropic and OpenAI, through proxies and heavily discounted token resellers, which they distilled and used for training data for their own models.

Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and OpenAI did), they distilled Anthropic's models to bypass the hardest parts of development. Chinese labs compressed 18 months of intensive research and development into just 6 months, and are now head-to-head with their American counterparts.

Anthropic tried to complain about this unauthorized "token theft", but they burned too much public goodwill with BS safety restrictions and users don't care. The US government is too busy fighting a war to help. Chinese labs are offering highly capable, cheap, open-weight models; exactly what users want. The community is happy to overlook any questionable methods Chinese labs used to build them.

The cope is incredible. There's people in this thread in denial that Moonshot AI is trained on exfiltrated Anthropic's model output, even when shown substantial evidence this has been happening since Kimi 2.X

Chinese labs were even paying an absurd $0.01 per Opus tool call trace, to get the quantity of training data needed.

Kimi K3 has reached the point of RSI, and no longer needs synthetic data generated by Anthropic/OpenAI models. K3 is now capable enough to generate, iterate, and improve its own training data recursively. The data exfiltration is complete.

We witnessed the most extensive industrial espionage campaign, probably ever, and nobody in the industry cares at all that it happened.

12 comments

I could perhaps get myself to care just the tiniest bit if the information that was supposedly stolen wasn't generated by "stealing" from everybody else. Either it is fair use to train AI models on whatever information you can get your hands on for everyone or for no one.
Stack overflow is pretending to be Claude now. I wonder if one can get it to say your question had already been asked.
Training data that’s still sitting in the pages of a book is not really all that useful.
that doesn't mean you get to steal it, and it means you don't get to complain when people reuse/steal it.
How dare they steal what we rightfully stole first!
Ask claude its name in Chinese and it says Qwen or Deepseek. Anthropic distilled Chinese tokens rather than create their own Chinese language training data.
Why should anyone care? I couldn't give a single fuck, in fact if what you assert is true (definitely not proven), I applaud Moonshot - seems like a very smart way to operate.
Shrug.

Hard to feel sorry for companies that created their empires by ignoring copyright themselves.

Also, 'most extensive industrial espionage campaign, probably ever' is absolute nonsense. They did not need to infiltrate the companies for this nor are you accusing them of stealing any trade secrets. This is only about whether they looked at their competitors' products from the outside (in the form of conversation tokens) and used it to improve their own product (by training). Hardly the crime of the century.

> 'most extensive industrial espionage campaign, probably ever' is absolute nonsense

You are completely underestimating the scale of what is happening here.

Chinese AI labs are actively facilitating an industrial-scale network of tens of thousands of bot accounts, that resell Claude tokens at 97% below official API prices. They buy subsidized Max 5x plans (sometimes with stolen credit cards), then split the subscription across dozens of clients and reselling the output. They are running a massive data-harvesting operation. Chinese labs and token resellers subsidize the cost of the tokens in exchange for the API metadata (detailed reasoning traces, model outputs, and tool calls) to use as high-quality training data for their own models.

They are buying Anthropic's own product, just to resell it below cost, just so they can capture the training data. Reportedly, they are paying as much as ~$0.01 per tool call.

https://x.com/yan5xu/status/2029743983522631698

I explained what is happening in this thread: https://news.ycombinator.com/item?id=48664814

You haven't explained how this is illegal or any more immoral than scraping the web for training data.

As you said yourself: They are buying the product. Then they are using it for their own purposes. That's more than Anthropic/OpenAI did for the open internet. That's more than Meta did when they obtained torrents of books in the early days, and then claimed that even though the data was obtained illegally they can still train on it just fine.

They paid for it! It's absurd to call this espionage!

> They paid for it!

They didn't though. The resellers are not buying via the official API, they're buying Max subscriptions (where tokens are priced ~10x below API cost), then splitting the subscription across dozens of clients and reselling the output as the regular API. Anthropic prices its subscription plans barely at cost, to bring in customers onto their enterprise plans where they can charge expensive API rates. Reselling these subsidized plans for price arbitrage is a TOS violation. It's not a legitimate purchase. Plus, a non-trivial amount of this volume is funded by stolen credit cards, so this "revenue" gets chargeback anyway.

The resellers then log all the the model output, then sell it to Chinese labs as training data.

> is it any more immoral than scraping the web for training data.

I think you'd acknowledge there's a difference between "We indexed public web pages" and "We deployed tens of thousands of fraudulent accounts to resell your subsidized plan for cheap, stealing your own customers, while collecting the data to build our own competing product" are very different actions. One can believe the first was wrong while acknowledging the second is far worse.

TOS violations are not espionage. Everybody who links up Claude to OpenCode is violating the TOS.

So from the largest industrial espionage in history we have left "They paid for the accounts but violated the TOS". And then you randomly add the claim they stole the money to pay for the accounts.

You have provided no evidence other than "Claude tokens are sold for cheap in China". As others have pointed out, that might also simply be counterfeit tokens generated by open weight models.

The western labs have established the precedent that all data they can buy beg borrow or steal is fair game. Turning around and crying foul when the Chinese labs follow their lead is hypocrisy.

Have a read of this detailed article, it's well sourced and documented that token resellers are logging the Claude outputs and selling them to Chinese labs. All your points are addressed in there https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...

I linked it earlier, but it seems you didn't see it.

re: labs purchasing model/tool output, see https://x.com/xkajon/status/2050445443889525235

re: model swapping, sure some providers may swap models, but there are many that don't, see https://www.hvoy.ai/en for a list.

Does anything about that strike you as particularly unfair, given the moral compass defined by Big AI?
The situation strikes me as morally ambiguous. The resellers are:

1) selling Anthropic's products at a 95% discount and redirecting Anthropic's own customers to themselves. A customer is far less inclined to buy directly from Anthropic when a reseller is offering an identical product for 10x less. This situation is highly similar to internet piracy.

2) keeping the token logs from Anthropic's products and selling them to competitors, so those competitors can build their own equivalent models. The resellers get paid per token log they deliver. This situation is highly similar to espionage.

May I ask you a personal question? What is motivating you to take up the frontier labs' cause in this way? Not a rhetorical question.

For my part, I'll happily disclose that I have an axe to grind. I think the major AI labs are an aggressive form of a cancer that's been ravaging our society. I want to see them fail, of course -- but more than that, I want to see the public develop an immune response to this.

I just can't wrap my head around why someone would expend so much effort speaking up on their behalf. They have, after all, highly compensated PR people doing that for them!

I suspect OP is the infamous Dario. We found his HN "anon" username.

That's the only explanation.

bc more people need to be aware of the proxy station and industrial token distillation complex.

many people i've replied to refuse to believe this is going on.

once you realize what's actually happening, and that you can get Chinese-lab-subsidized tokens at a >95% discount, why would you ever pay full price for overpriced APIs?

if Ford bought hundreds of millions of dollars worth of Hyundais, put extra instrumentation in them, and resold them at a discount to customers who agreed to the instrumentation in exchange for the discount, is Ford doing industrial espionage?
You skipped the part where Ford buys the cars at 90% off and sells them at 80% off, at a profit. Then gets paid by competitors for the driving data.

At the same time, Volvo is running the exact same hustle, except they buy the cars with stolen credit cards, so they get the cars for free.

You skipped the part where Hyundai chose to sell the cars at loss in the hopes of eventually gaining a monopoly position.
Misleading on both counts:

1) Anthropic tokens via subscription aren't sold at a loss, they're sold at cost.

2) Subscription plans are not sold in hopes of eventually gaining a monopoly position. They act as a loss leader designed to get a foot-in-the-door and funnel companies into costly enterprise plans, where Anthropic can charge full API rates.

Don’t sell your cars at a loss then.
Where are you getting this stolen credit card thing? That's so random.
Someone was claiming that it's not an issue for these proxy networks to create thousands of bot accounts and resell Claude's output because "they are buying the product" and "They pay for it!", which is a nonsensical position.

I responded that these resellers don't always acquire these accounts legitimately. They often use stolen credit cards, educational discounts, or resold compute credits to acquire them at essentially zero cost. They're not always paying customers.

That's one reason token resellers are able to price so cheaply, they acquire the goods for free.

Anthropic and OpenAI eat the loss.

>The community is happy to overlook any questionable methods by Chinese Labs.

Using the US-based models are arguably even more questionable. You have to be content with the OpenAI and Anthropic literally scraping the entire internet. They've all pirated content, scraped against ToS, ignored robots.txt, bypassed paywalls all to train their models. It's well known these AI labs have ingested the entirety of Annas-Archive into their models, the largest collection of books ever assembled.

They didn't credit or compensate literally any artist, author, scientist or publicist in the creation of their models.

They've lobbied against local Governments to shove big and loud datacentres in peoples backyard. They've polluted local water supplies, they've doubled energy costs for these regions. They've tormented locals with subsonic frequencies.

They've given access to the DoD to use their models to kill people, or assist in killing people. They've used these models to enable mass surveillance, allegedly not domesicially but since when can we trust any of the 3-letter agencies.

Using "US Models" is not the moral high ground you think it is. Kimi saying its Claude, Gemini or ChatGPT is not the "substantial evidence" you think either.

>We witnessed the most extensive industrial espionage campaign, probably ever, and nobody in the industry cares at all that it happened.

Because they stole for every one of us without permission. Thousands of my comments on this site and others (Stackoverflow, etc) are all used in their training data.

Not to mention, OpenAI has allegedly just stole tons of internal Apple documents... I guess we'll just ignore that too.

assuming the k3 model weights do indeed get published, if your model of the world is "achieving RSI is beneficial and K3 has done so," this feels structurally different from ordinary industrial espionage, because the knowledge has enriched the commons

more like silk than capacitors

if, again, your model is that RSI will be beneficial, why wouldn't making it available to all unlock more benefit globally than not doing that

Sorry, but I have a thought that's off topic: If Kimi is good enough to improve itself, what's preventing someone who owns a big datacenter and nothing else, to just run Kimi to do AI research, thereby rendering the frontier labs more or less unnecessary?
Not at all curious you keep citing sources only from a platform run by an incompetent rascist arse wipe that you constantly boost all over here.

"The community is happy to overlook any questionable methods Chinese labs used to build them."

Oh and you aren't willing to overlook all the actual illegal activities, as in actual court cases that the US labs and corps used to train their models?

Industrial sabotage? Pa-lease! get off your pretend high horse and stop making us laugh.

> We witnessed the most extensive industrial espionage campaign, probably ever

This is the funniest way of saying “going to a company’s website” I have seen in my entire life

>through proxies and heavily discounted token resellers

Could you explain a little more about how this works? Are you saying that the Chinese run or have backdoored something like OpenRouter?

Have a look at https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens... and https://x.com/yan5xu/status/2029743983522631698

Chinese resellers acquire hundreds of Claude Max 5x accounts and set up a custom proxy server. Customers point their ANTHROPIC_API_KEY at that proxy, and requests are routed to Anthropic through one of those hundreds of accounts. Because one $200 Claude Max 5x account gets the equivalent of ~$2000 in of API credits, these resellers can resell Anthropic tokens at a massive discount, undercutting official API prices by more than 90%.

To cut costs even further, these accounts are funded using educational discounts, startup credits, or stolen credit cards.

The resellers log all data traveling through their proxy networks, which they then resell to Chinese labs as high-quality training data for significant profit. https://x.com/xkajon/status/2050445443889525235

The resellers also loan these proxy networks to Chinese labs, allowing them to can run distillation attacks on Anthropic, while blending in with regular user traffic. https://www.anthropic.com/news/detecting-and-preventing-dist...

This is a widespread tactic, there's hundreds of proxy resellers operating. Some even offer enterprise SLAs.

oh wow, an entire seedy underbelly I was unaware of. Thanks, great reply. Appreciated!
> nobody cares at all that it happened.

Who in their right mind would care? Why care? Misplaced patriotism?

"A thief who steals from a thief has 100 years of forgiveness". Spanish proverb.

In fact, I would be very concerned about the sanity of someone who cared about this sort of thing, unless they were Dario themselves.

> nobody cares at all that it happened

Oh, no. I wouldn’t say that. If that happened, I definitely care: I’m positively delighted about it.

They stole from me first. And are spitting in my face and telling me they’ll take my job while they do it. I have negative sympathy for them.