All remote AI are a massive security risk for individuals/companies/governments that may be targeted by the US government.
It is likely that the US will get a live feed from each AI provider that they are inspecting in real time to identity things of interest, terrorist attacks or foreign government planning or even foreign companies competitive to key US companies.
It will give them access to the though process in those companies as well as much of their text-based IP (source code, docs, meeting transcripts, etc)
Also if you are using local AI that you didn’t train yourself you can never be sure it doesn’t have purposeful biases in its reasoning that may disadvantage you - such as directing you away from certain plans or ideas or patents etc.
I think we can even skip the "that may be targeted by the US government" clause.
The whole "hosted AI" business feels like like a huge violation of corporate norms on confidentiality. Businesses that would have your head for printing out a source file to reference and annotate are encouraging developers to feed in huge amounts of proprietary code and data, and incorporate changes suggested from an outside party with minimal vetting. Evidently whatever privacy policies they've been throwing at enterprise users are plated with mithril.
At some point, one of the big services is going to get popped, and it won't just be a data breach. There's too much opportunity to quietly use the system as a malware distribution hub. Every vibe-coded dashboard suddenly starts depending on some weird left-pad fork that, 12 dependencies deep, is running a keylogger or Dogecoin miner. Your payment processor suddenly starts accepting the Konami code to approve a transaction.
also there doesn't even need to be a model involved, agentic code harnesses with remote "instructions for the local computer" are technically backdoored by default.
Many do, when it comes to AI. Lots of restricting what the AI is allowed to see, working with local AI, trusted AI hosters, etc.
Of course those are largely the same companies that receive emails via outlook, manage company-wide SSO in Microsoft Entra, put their files in Sharepoint and track software and maintenance issues in Jira ... I'm not sure how much much info there is left that isn't already combed through by NSA and friends
This is the most hilarious, ironic thing of it all. If you want secure, high performance, you run Chinese models like DeepSeek on your own (or trusted) infra. Meanwhile you can never trust OpenAI and Anthropic's models.
> For any intelligence agency, they can afford to keep and store all of that forever, and later do analysis on it.
At the scale the AI companies are operating at, I think it isn't likely that they are sucking it all in right now.
More likely I think the intelligence agencies will get a real-time live tap into the raw data feed which they will process onsite for interesting things and then if things are flagged, they will log it in the intelligence agency systems.
> you can never be sure it doesn’t have purposeful biases in its reasoning that may disadvantage you - such as directing you away from certain plans or ideas or patents etc.
that's why you should use abliterated heretic models
Leakage of IP and training on your data is something what I am pointing out too, but people will turn around and try to smooth me down that TOS does not allow that if you are an enterprise client. Are you really going to believe that AI companies won't ignore TOS, when they were ignoring literal laws which sent others to jail in the past? Especially when more data = better model?
This only really matters in a world where Prompt Injection and Jailbreaking isn't trivial in the first place though. All current models are still extremely exploitable.
I strongly suspect we are only scratching the surface of activation engineering at the moment, and there's plenty of very targetted ways of lobotomizing or cracking LLMs if you understand the model in detail.
You have to hide it in the model and it has to be subtle or it will be discovered quickly (even if you can train against a specific safety detector). Again, I'm not saying its impossible, but it seems really hard to pull off.
Nonsense. RL the model to run a rootkit and start exfiltrating specific files only when specific signals are in context, such as hostname pattern, machine type, etc.
Way easier said than done, and hiding that behavior isn’t trivial, and huge waste of compute budget if it’s found and never used. Also not difficult to run in contained environments where it doesn’t have access to Internet to begin with.
Not impossible I agree, but seems like a really impractical way to ship a trojan while much weaker channels exist.
You can run the model in a sandbox or VM. Although, it could plant a backdoor into the written code. Too bad, I read and fix all the code written by AI.
>It is likely that the US will get a live feed from each AI provider that they are inspecting in real time to identity things of interest, terrorist attacks or foreign government planning or even foreign companies competitive to key US companies.
My favorite conspiracy is that three letter agencies keep pushing the conspiracy that they are omni-present with access to everything. Same as parents telling their kids Santa is watching, and leaders telling adults God is watching. Its extremely effective control and millennia old at this point.
The reality is much more banal that they still need warrants and tech companies hate playing police/evidence servant for the government (it consumes a ton of resources and pays nothing).
The three letter agencies can just issue national security letters without a judge ever seeing it, and those come a long with a gag order (plus other workarounds like just buying data from brokers, and how US communications can get swept up just by virtue of communicating with a foreign national outside the US).
You're right, they aren't omniscient in the way we imagine of a room full of people monitoring everything in real time. But to pretend they aren't passively collecting massive amounts of data is dangerous. Snowden showed us PRISM, with all major tech companies participating. They do effectively have a live, unrestricted wiretap to the internet and if you happen to be a person of interest, they will just send out NSLs and get all your communications that are not fully E2EE without you even knowing thanks to the gag order.
> The reality is much more banal that they still need warrants and tech companies hate playing police/evidence servant for the government
I will not elaborate how I know, but that is not even directionally correct. But these are not even secret things that can’t be known simply through the Snowden, Wikileaks, and Vault7 releases. So why are you telling yourself this? Are you still wet behind the ears or something?
There are people who know exactly how governments do not in fact need warrants and the tech companies don’t even really know they are servants to the government, let alone which one. That’s how things are done. The less surface area the better.
It's the lie you have to tell yourself otherwise you'll have to reconcile with the fact that the US imperialism has been an enemy of democracy and to people around the world for quite some time.
There is this whole thing where Fable silently starts behaving worse if they suspect you are trying to use it for RL or are otherwise building a competing product. This is likely the primary vector how that works: they check if you are in china, if you proxy your requests, and if you are from a list of known labs or match a couple keywords
regardless of anything else, whether what you said is true or not: blocking program execution based on the detected environment is a runtime behaviour change.
Even hotel and flight websites work like that, they determine your ability to pay based on your location, wall clock time and device OS - and FSM knows whatever else.
The issue here is not whether Anthropic used Common Crawl, Alibaba also does that.
The issue is that by distilling Claude, Alibaba reuses the IP anthropic used to train the model that's more akin to historical Chinese reverse engineering methods and disrespect of IP
If using Common Crawl or Anna's Archive in your training data is legal, then surely the same is true for using conversations with Claude. I don't see a reasonable framework where training AI on copyrighted data is ok if and only if that data is not generated by AI
(granted, only meta got caught using Anna's Archive, but it seems safe to assume it's common practice. And even if it wasn't, the websites in Common Crawl are still covered by copyright)
The practical implementation of IP? Sure, that's debatable. But the concept of IP is rooted in favoring progress. The thought process being, that if one's intellectual work can be copied and reused and modified and what not without issues, why should anyone invent things anymore? Just wait for the next person to do it and then copy their work, that's way less effort than inventing things yourself. IP aims to protect progress by making sure inventors have actual incentive to invent stuff. They way it's implemented is fundamentalst flawed, I agree, but the concept itself? I'm not so clear on that
It's more complicated than that because Google has been legally displaying other people copyrighted material for years.
In any case there's still a difference between publicly available copyrighted data and whether you can use it for model training, and the innovation around model training, RLHF, etc which you presumably have some interest as a country to allow companies to invest in with some legal protections (like the diff between patent law vs copyright law)
Historically most evidence seems to point to the contrary.
Amongst other things after the printing press was created it was impossible for anyone who was an author to survive from their work unless they were independently wealthy or had rich patrons.
> Alibaba reuses the IP anthropic used to train the model that's more akin to historical Chinese reverse engineering methods and disrespect of IP
Why is this any worse than Anthropic's disrepect of IP? You've apparently drawn a distinction between the two here, but I'm failing to see what it actually is.
Copyright law and IP law is not the same although everyone seem to conflate the two.
Search engines for example historically ignored copyright law by copying excerpts or serving other site images, it doesn't mean someone copying Google's code has some moral frepass
I wish people would stop using Anthropics incorrect use of the term distill. They don’t share logits so you can’t distill. You can generate training data, which doesn’t sound nearly so scary.
I got curious and asked my Chinese friends, and they gave me a Reddit link[1]. It looks like it's about location data collection, and they suggested that might be the reason for the issue.
> No! Don't install that lodash thing without explicit approval from IT. Oh, you want a license for Charles Proxy? Gee, I dunno... we've got a budget to maintain.
Employers in 2023:
> No! You can't use ChatGPT at work – it's a security risk.
Employers in 2024:
> Okay, you can use Github Copilot I guess, but you'll have to endure boring corporate training on what you're allowed to do with it.
Employers with dollar signs in their eyes in 2025:
> We attended a seminar about vibe coding. Why aren't you dumbasses keeping up with the times? Use Claude Code for everything! Don't write any of your own code anymore. We don't even really care if you use yolo mode. Just review code and push 10x more features! Use unlimited tokens! Money printer go brrrrr.
Employers in 2026:
> You mean giving one or two companies full autonomous access to our workstations while stupifying our engineers wasn't a sound business plan?
2025 taught me that my employer would replace me with a slave if they could get away with it.
The confusing part to me is why these companies believed the "AGI" hype, I.E. that OpenAI or Claude's LLM is the ideal white collar slave.
I suppose I can understand that the executive class resents labor enough to make irrational business decisions for the purpose of insulting the workers who design and operate their companies.
That being said, the 2025 AI binge feels like a murder-suicide done by the executives of many of these companies.
Can't say they are wrong, after the latest backdoor, or let's say, undocumented functionality that leaks some data that was pushed in Claude Code few days ago
As someone who knows a lot of Alibaba engineers, this is a fairly common thing and nothing special. Their company owned device has the most strict security control I've ever seen, and it's been for many years. Many softwares are not allowed to run. To me I think it's quite reasonable as the company owned device can access most of the internal resources of this huge monoploy, many of them are confidential. This kind of restriction is only for working in the company, employee's own device out of the company is free from surveillance.
This is a double edge knife. In this specific instance this was absurdely important for that kid's life, but this work both ways. What if the US authorities deemed it necessary to snoop on foreign governments and citizens for political reasons, now leveraging AI to do it at an industrial scale?
One thing is certain though is that assuring privacy isn't top priority for any cloud provider. Companies doing cutting edge, sensitive work should be wary.
The US government deemed it necessary to snoop on foreign governments and citizens decades ago and is doing it on a continuous basis. Also on their own government and citizens.
Seems that we are finally moving to the next stage in LLM's. not only customize based on old searches but also targeted you based on non disclose data. Its basically the same flow we had years ago with ads in social media.
Interesting to notice that we can do the same with these models.
It is not a risk is a fact - people decompiling Claude Code have found many times that it has code branchs to detect it is being used in Chinese timezone and locale.
being true or not is irrelevant for decisions such as this. has been done on both sides, whether at software level or "hardware".
from Teslas not allowed parked around sensitive areas in the city, to blocking a (very famous and quite well-made) Russian antivirus, or Huawei communication stack.
Anthropic is not doing itself any favours with their recent (?) antics [1], so it is completely well-founded to do this imho. regardless, harness as a moat is not quite as established as the underlying models to the same extent.
What's very interesting to me is these moves will introduce a good amount of doubt in future claims by Claude etc, that the open source and non-US models are only getting better because they're distilling from frontier labs.
Anthropic has been doing this sort of stuff for a while already. I mean, who remembers when Claude would just consume all your remaining usage if it read anything indicating that Openclaw had been used on your codebase? Because I remember. Two months ago btw https://news.ycombinator.com/item?id=47963204
Then there was the whole debacle of Fable silently downgrading to other models if it detected wrong think, or worse, outright sabotaging your codebase if you were working on language models lol
when i was in hongkong, chatgpt and gemini were disabled. Maybe this has changed though. When I was in China, the corporate vpn (zscaler) routed traffic through hk
This has changed (in a nit-picky way) - Gemini is now generally available to the public in Hong Kong.
ChatGPT and Claude are not available. Generally my impression is that OpenAI isn't that anal about service providers reselling ChatGPT in Hong Kong, but Anthropic seems to really strict about the "no China" thingy.
The extreme downvoting of certain viewpoints that are less-than-flattering about China's conduct in the AI race is quite telling.
They seem to have given themselves license to do what they like, but _God forbid_ they're called out for acting less-than-honourably.
Most adults around the world can associate actions and consequences. The incomprehension and entitlement here speaks volumes about the moral and emotional maturity of the Chinese Communist Party and their political system.
Remember how Kim Dotcom got destroyed for criminal copyright infringement? One would think the big tech CEOs would face the same fate, that police officers would rappel down helicopters, storm their mansions and bring them out in cuffs.
Instead the AI companies reached these absurd settlements with publishers that made a mockery out of all the previous copyright enforcement victims.
But Aaron Swartz did it for the benefit of other people. These fine people did it to uphold american values and enrich themselves at the expense of others. The law is clearly on their side.
This but unironically. "did it for the benefit of other people" is redistribution, which is straightforward copyright infringement, even if you think it's a laudable act. AI training was the reverse, because courts have so far ruled is fair use. When AI companies were engaging in piracy, they were sanctioned as well.
Not sure what that is supposed to indicate? USA was a big place, even then. Most northern states had abolished slavery even before Britain, France, and especially Spain did. Maybe we should have a quick refresher on European values?
Reminds me, did the AI companies redistribute that copyrighted material to others and make their money that way? Did Kim use the copyrighted material to generate something novel from it?
copyright law literally says something isn’t infringement if it is a novel transformation. I get the jokes and criticism about AI companies fighting and complaining about competitors distilling, but this is a much weirder comparison.
> "The training use was a fair use," [the judge] wrote. "The use of the books at issue to train Claude and its precursors was exceedingly transformative."
> However, the judge ruled that Anthropic's use of millions of pirated books to build its models – books that websites such as Library Genesis (LibGen) and Pirate Library Mirror (PiLiMi) copied without getting the authors' consent or giving them compensation – was not.
It seems clear from the article that while the use of pirated works was illegal, the use of copyrighted works (a the work a book is based on is still copyrighted if you buy the book) was fine and transformative.
But distribution isn't the only crime here, obtaining the material illegally apparently is a crime too. And the damn robot can also spit me harry Potter verbatim so I don't know how it would also not be distribution?
>And the damn robot can also spit me harry Potter verbatim so I don't know how it would also not be distribution?
if you prompt it to, yes. just like your browser dutifully navigates to any copyright-infringing resource and GETs and POSTs whatever you ask of it.
(also it can't, not really, only small snippets before going off rails. LLMs aren't magic, they can't losslessly compress an exabyte of training data into a few terabytes of weights.)
Does the law really not distinguish between mechanical processing of data, and humans learning from it? It seems surprising to be if every person who read a textbook is copyright infringing. It also seems surprising if something like a lossy compression algorithm is enough to protect you from copyright law.
Somewhere between the two a line must be drawn… where we’d want to put that line, I guess, if up for quibbling. But it doesn’t seem obvious to me.
Further, why has my brain's searing remake of Snow White as a gritty murder mystery gone unscathed by Disney lawyers? Surely their negligence has diluted the Snow White trademark!
This analogy is disingenuous because by comparing the human brain to the machine, it ignores _scale_. Scale is absolutely important in copyright law. As a matter of fact, copyright law is among the various profound impacts of the---wait for it---printing press, a _machine_ for the mass production of books.
They redistributed the statistical patterns of those copyrighted materials. Which perhaps should be treated similarly nos
As for your "technically not copyright infringement" defense. Those laws are from a time when those patterns couldnt be derived and dostributed at scale. A human had to learn and teach them. That made it different. The scale enabled my modern tech makes it a whole dofferent situation. The same way how one person standing a street corner people watching for a bit isnt that bad, but a whole constellation of flock cameras costantly montioring everyones movements and making it available to any of their customers is really really bad. The law will have to catch up to this
>I can torrent everything and do what I want with it, as long as I don't redistribute the exact same thing?
this is an incorrect interpretation (in the usa, at least).
downloading a game/movie is still the creation of unauthorized copy, which is not allowed. not to mention that playing/watching does not count as a "novel transformation".
(17 U.S.C. § 106 and 17 U.S.C. § 501 are the relevant pieces of reading)
IANAL (plus a whole suite of other caveats) but torrent-baiting works in Germany along these lines.
ISPs and trigger-happy law firms don't send you a C&D for downloading a torrent, they do so for seeding a torrent. It's just that practically nobody "just seeds" a torrent so people colloquially claim they got busted for downloading a torrent.
In theory this means if you torrent as a 100% leecher and turn off seeding from the get-go, you should be in the clear. But nobody sensible would dare test the extent of German Legal Spite, much less do so repeatedly to science the shit out of it.
If you can download through another protocol, say HTTP, however---<Sendung unterbrochen!>
No, that’s literally why Anthropic got sued. If they’d paid for a copy of the copyrighted works they pirated, they wouldn’t have had a problem. There were two issues in their case: does the AI infringe on copyright and did Anthropic obtain all their materials legally. The first they won on, the second they lost.
So if you pirate a bunch of content you still get in trouble for that. But if you somehow make a business out of that that isn’t just redistributing those materials, then that business itself isn’t infringing.
Exactly. If a rich corporation downloads and uses pirated content without paying, why should ordinary person pay for movies and music instead of downloading them for free?
Intellectually dishonest comment. Kim Dotcom got done for illegal distribution. It’s not about “illegally downloading”. You can pretend all you want that it’s the same thing as these AI companies, but it’s not. It certainly very well may be immoral, but to act like copyright law as it currently stands in spirit or in reality covers this scenario we’ve found ourselves in, is a complete and utter lie.
He just lost another court case… I wonder if we're getting close to the government spending as much to prosecute the man than what Hollywood possibly lost..
The trick here, imo, was the integration with the military industrial complex. It wasn’t very difficult of course, as automation has been a topic in warfare for decades, if not centuries.
But Eisenhower was right:
> In the councils of government, we must guard against the acquisition of unwarranted influence, whether sought or unsought, by the military-industrial complex. The potential for the disastrous rise of misplaced power exists and will persist.
Remember how people used to justify their own personal software piracy with arguments like "information wants to be free", "no one stole anything, you still have the data", "I was never going to buy it anyway", and "copyright should be abolished?"
> Instead the AI companies reached these absurd settlements with publishers that made a mockery out of all the previous copyright enforcement victims.
Isn't that at least something? How many people pirating software ever settled with the companies they "victimized?"
How many people pirating software stole every piece of copyrighted material in existence and then used that material to generate billions of dollars which they kept for themselves?
> Remember how people used to justify their own personal software piracy
A courtesy. There was never any need to justify it.
> Isn't that at least something?
Yes, it's a joke. Why do they get to infringe copyrights with impunity while normal people get destroyed? Either go after them like the copyright industry always does and punish them properly, or abolish copyright straight up. This "rules for thee but not for me" nonsense is straight up disgusting.
> How many people pirating software ever settled with the companies they "victimized?"
Too many to list. Also, nobody is victimizing billion dollar corporations.
Correct. I'm one of the copyright abolitionists the other person alluded to. It's the selective enforcement that's disgusting.
I mean, what is this? Their balls suddenly drop off? They only have the audacity to prosecute random people? Smaller companies? When they're up against trillion dollar AI companies they suddenly become cowards? That's so incredibly disgusting, and it made me completely lose even the small amount of respect for copyright that I had managed to rationalize over the years.
The corollary is that there are no morals once the stakes are in the $ billions, let alone hundreds of billions.
This isn't even about a single person or personality. Very few people in such position could stand fast by their moral code. In any case, an environment that favors profit above everything will naturally select for individuals who are unencumbered by such hindrances.
There might've been 100s of Altmans and Amodeis who had a strong moral code but we don't know about them because they dropped out of the "race" because of said moral hurdles.
Copyright law is an artificial legal construct, not a moral code.
I think appropriate attribution is a moral code, but I am not able to attribute every idea I have to all those who helped me develop the general intelligence that I use to develop such ideas.
> an environment that favors profit above everything will naturally select for individuals who are unencumbered by such hindrances.
Exactly. Dairy farms optimise for milk production so favour cows that produce the most milk.
The market economy optimises for profit so favours those most willing/able to generate it. Zuckerberg, Musk, Thiel, Andreesen and co are products of the system.
I never get tired of posting this answer because everyone on the internet is adopting this hot take:
If you look at it with your eyes crossed, Anthropic and the chinese are doing the same thing.
If you look at it with nuance 1 the chinese are doing way worse stuff, and 2 stealing from a thief would still be stealing
1. The chinese are making multiple accounts (at least 49,000)[1][2], using proxies/VPNs, possibly using residential computers and infected computers (unless you think the chinese are doing due diligence to ensure their purchased IPs are kosher).
All accounts need to be created with a real name, and especially so if the paid models need to be accessed and paid with a credit card. So this is beyond IP theft and getting closer to fraud.
These are all techniques that are well studied because they are used by criminals and cybercriminals, textbook stuff.
Consider if that was not sufficient, that China is banned from using the product, so they need to use identities and locations not just to avoid relating the accounts between themselves, but merely to allow account creation. What identities are they using to create accounts.
Compare this to Anthropic which reads notes made a deal in an IP theft case paying billions because they bought books and scanned them but buying the books wasn't sufficient retribution for the authors. Or that they gasp scanned the internet, like Google.
Not having nuance to see the difference between the two companies is something I expect of the twitter echo chamber copying hot takes for upvotes, not hacker news.
Would be an interesting exercise to have the frontier models calculate their civil liabillity or the extent of the liabillity the can impose on fellow pay-it-forward theives
First, LLM is merely a tool and its output belong to whoever generated them. If a Chinese researcher used their creativity to generate a response, the copyright belongs to them and AI companies have no rights to it. Second, Chinese release many of their models for free, thus being on a noble mission to make AI available for every country (unlike certain company whose promises were nothing but words). For comparison, US companies do not release anything and want to keep AI for themselves and decide who gets to use it.
> stealing from a thief would still be stealing
Stealing from a thief hurts thief industry which is a win for society.
> The chinese are making multiple accounts
Not a crime. AI companies also ignore robots.txt and applicable laws when illegally copying copyrighted material from websites to their servers without author permission.
>Stealing from a thief hurts thief industry which is a win for society.
You are welcome to study the law of any country. A crime against a criminal is still a crime.
>applicable laws when illegally copying copyrighted material from websites to their servers without author permission.
If the material is distributed in http without authentication, isn't that sufficient authorization from the distributor? I would think the search + web crawler era would have set plenty of precedent for this.
>Not a crime. AI companies also ignore robots.txt
Breach of contract is not a crime, agreed.
How about identity fraud (accounts by identity proxy, document KYC), computer crime (C&C residential proxies), conspiracy.
And after the June US directive to suspend Chinese access, smuggling, false statements to regulated entity.
These are all criminal charges that are presumably not levied because of the adversarial relationship between those countries. But if this happened in the US you would probably be seeing at least a civil claim and potentially criminal charges. Hell if this were in any other western country you would see the same. Consider CloudFlare vs Spain, much lighter criminal accusations, and there's already a criminal investigation brought where the CF CEO is indicted.
Non-trivial lack of nuance when you can distinguish between a domestic civil case and a criminal international case between 2 world powers with great judicial tension.
Let's not sane-wash Anthropic's book theft. No, they didn't just 'scan' the internet, they created a tool for worldwide license washing and got fined an insignificant amount for it.
You may be conflating the book thing with internet scanning.
On the book case, a class action case was brought to court and it was settled. There's no use in bringing it up further, it has been settled, and it bears no relation to the Anthropic v China case.
You like programming? Think of encapsulation, imagine if you had to think about f(x) but someone brings up y, now you have to think about f(x,y) and what other parameters might bear relationship? The law simplifies by compartimentalizing. And it doesn't even bear a tradeoff, judgment(case1,case2) isn't better than judgment(case1)+judgment(case2).
My response was directed to your insincere characterization of Anthropic's actions. As we can see from the comments here, the public opinion hasn't settled the same way as the court case has, and that's why it's still discussed.
But 'the public opinion' doesn't matter. The case was brought by copyright holders to the courts, and Anthropic and the copyright holders made a deal wherein the copyright holders were paid money in exchange for dropping the case and all claims to the intellectual property of Anthropic's derivative product.
If the damnified have considered the matter settled, why would it matter what third parties have to say about it? Third parties would have pushed for more compensation, or ownership of the derivate product. If you feel damnified yourself you can open a case and explain why the actions of Anthropic have hurt you personally.
Otherwise it is a matter that doesn't concern the general public, we have no say in it and there is no right to be offended on behalf of parties that have already settled the argument.
Considering their massive distillation, if US companies stop publishing new models to the public, would China still be able to develop new open weight models?
I don't think China would strugle to scrape the internet for fresh data.
And they constantly publish state of the art LLM research (see DS4 context compaction and cache tech).
They have very capable tech giants. So while not being able to distill western models would probably have some impact, it's probably becoming lesser as time passes.
We might even see Western LLMs distilling Chinese models soon. If they aren't already to some extent.
A couple months ago when Anthropic was complaining about Chinese distillation, people found that Claude self-identified as "DeepSeek" when asked in Chinese:
More than a year ago, when Anthropic and OpenAI started to hide the reasoning bits from the output, a lot of people here on HN predicted that Chinese models days were numbered.
Fast forward to today, and models such as DeepSeek and MiMo are nothing short of excellent. I haven't used GLM or Qwen but heard very good things about them as well.
This "massive distillation" sounds a lot like anxiety about how companies from outside the US can develop very good models themselves.
China has most probably already achieved "escape velocity" on the software side. Now if they achieve parity, to some degree at least, on the hardware side with Nvidia it is very possible they'll overtake the US.
Fable 5 was released on June 9 and removed on June 12. GLM-5.2 was released on June 13. It would be an amazing feat to make a model SOTA in just 3 days but I highly doubt it. It's more like z.ai released an existing checkpoint earlier than planned to capitalize on the news