Hacker News new | ask | show | jobs
by rustyhancock 1 day ago
He's also completely missed what many of us see as the primary the issue, HF was attacked by a US frontier model.

HF could not be helped by US frontier models because of the "safety" features they have.

HF had to use an open model from china.

Anthropic wants to add those "safety" features to open models - especially from china.

End result would be HF hack would have continued atleast until the Monday that OpenAI engineers finally walked back into work.

4 comments

I'm very confident that it was staged and coordinated. This propaganda started because they cannot evolve their models further. See Fable and Sol, they are lame. They seem incredible at first but the more you use you can see the trickery. It is a matter of time for someone to prove they are marginally better only because they inject more information to the harness at server side.
I’m also into this conspiracy theory: ai created bug, ai hacked it, ai fixed it.
just like ai will create amazing tests suits for the broken code it writes, if one isn't carefully watching...

EDIT: of course it probably helps to have an up front tdd test suite, but often isn't the case.

Live overflow released the video today about it. It is very compelling but I guess there are holes on their hypothesis as well.

The summary is that an instance was running on certain benchmarks without any limits or supervision and the AI decided to cheat by exploiting a silly series of vunlns.

The public narrative around conspiracy theories is baffling for sure: what rational basis is there for assuming them wrong? Often none.

Incentive and ability are what should be looked at. There, things get far more interesting: what is the current state of AI employed by the US intelligence agencies and what do they use it for?

Having the public convinced, their "superiors" would only do everything in their best interest, even without anybody knowing for sure, is Huxley's Brave New World in real life.

Have you seen any of the emails Mr alt-tab-man sent regarding his enterprises? That guy is worse than fetid trash
that's some real anti-truth you've got going on there. because we shouldn't assume that others are acting in our best interest, that any random assertion that there is collusion against us should be by default accepted as the truth.

shouldn't any reasonable person when presented with a line of thought that has no substantiation at all conclude that they just can't reasonably be expected to support or reject the theory?

? That's total nonsense.

Positions of power attract people who want to use that power for their own gain. They abuse it regularly, whether it's politics, enterprises or even charity.

To assume, that somehow wouldn't happen when you don't look is beyond absurd. Of course it does.

So I pointed explicitly at US intelligence agencies, because there corruption is rampant and oversight virtually non-existent, respectively itself involved in the corruption.

What you suppose there is a grotesque "don't ask don't tell"-complicity.

In reality, trust has to be earned and that deservingness has to be verified.

Everything in this space is manufactured, the plebs is fed very deliberate information for manipulation. I don't think there is a debate about it. All the PR stuff is marketing.
I case you're wondering, HF = Hugging Face
Every time I read Hugging Face, my brain first jumps to facehugger from Alien. I really wish they chose a different name...
Actually I wrote something similar yesterday as a reply to the (since flagged) comment below. Glad to see that more people have this association. But apparently the founders of Hugging Face didn't have it, or else they probably would have chosen some other "friendly" emoji for their startup (which initially developed a chatbot for teenagers - that makes the choice of name feel even creepier if you ask me).
Same here. The creature seems more fitting than the emoji too.
Worth checking, elsewhere in this thread someone points out glm only assessed the damage after the fact, it didn't stop the attack, and Hf apprently never sought access to a trusted defender program with a closed model either the open model saved them farming might not hold up.
They wrote themselves that before the incident they were denied from the trust program...
I think the fair nuance here is, an administration which used Executive Orders to force guardrails, meaning US companies must retain them, even if they might want to drop them now.

And on top of that, with low/no guardrails, people call you a child pornographer(grok), so the public is also against it. Yet mysteriously few complain about Chinese open models being child pornographers.

So even if your goal isn't ethical, but just fiscal, it's reasonable to say there are two standards. And to complaint in some way.

I don't think banning is going to work, that's just silly. And over the next few years, everyone and their dog will have local GPU compute to train locally. People have home labs, the bar isn't that high, and eventually large text datasets will escape from Anthropic and other companies, allowing for comparable training.

It's a genie that's not going back in the bottle, the bottle is smashed.

The only reasonable outcome would be section 230 style carveouts so that there is zero liability for anything a model does.

Because having guardrails on corporate models barely months ahead of open ones, which will never be restricted, is entirely pointless.

>child pornographer(grok)

People call grok that because deviants were abusing grok's ability to edit images and post them publicly on X to strip people - including children - of their clothes, from their public photos. Then when there was backlash, Elon laughed it off. It took half the world opening investigations against X for violations of existing regulations for action to be taken.

The leniency that internet companies get in terms of dealing with illegal content comes with the expectation that they're making reasonable efforts to control the distribution of said content. X, and Grok, were actively supporting the production and distribution of the content in public.

It is very different from someone creating such images in a private account, and definitely very different from someone using a local model to do it.

My entire post was how it is unreasonable to have a dual standard, and my point was it's really irrelevant if it's a model you download and use locally, or if it's a model hosted remotely, or hosted and created remotely. You're not really providing any sensible reason where the line is, except "public company", which is, again, the entire point I'm making.

The only realistic, non-double standard is that the creator of the model should be 100% responsible. What on earth does it have to do with who's hosting it?

And by this metric, aren't all the uncensored models on huggingface, child pornographers? And if so, why not? Provide tangible, real, sensible reasons please, and after all, isn't hugging face a company?

You know, people are all over the place on this. I see people complaining about guardrails, then in the next breath complaining there aren't enough. Complaining that open models are the thing, but then creating double standards.

So once again, what is your actual reason why it's different?

It isn't about "public company", the line is "posted to a public social media platform, with the owner indicating amusement".

If someone uses something one of OpenAI's hosted models to generate illegal content, it isn't automatically spread to anyone that clicks on the post the original image is from. On top of that, OpenAI does not explicitly endorse the use of their models for such activities the way Musk did.

If someone uses a local model to generate illegal content, same thing applies.

I have a hard time believing that you're arguing in good faith. The differences are so blatantly obvious and they don't even require any discussion about guardrails. Grok was being used to edit images and post them on public social media. The technology to edit images has existed for decades now. So here's a direct analogy via Photoshop:

1. In the case of a local model: Someone using a local image editor to edit photos into illicit material and manually distributing them.

2. In the case of a hosted model: Someone using a cloud image editor to edit photos into illicit material and manually distributing them.

3. In the case of Grok: Someone using a cloud image editor to edit photos into illicit material, the cloud image editor automatically distributing it to the public, and the owner of the editor endorsing the material as an exemplar of what their editor is capable of.

In 1 and 2, the editor (ie the AI) was just a neutral tool. In 3 the issue is that the company explicitly endorsed and distributed the material, no longer neutral. This is why it's reputation is of being a CP machine.

The leniency that internet companies get in terms of dealing with illegal content comes with the expectation that they're making reasonable efforts to control the distribution of said content.

Your words. Grok was an example, the owner is irrelevant in the discussion.

Really showing that you aren't engaging in good faith and your repeated incorrect interpretations of my words are out of a desire to cover for Musk.

To spell it out for you further, the owner's actions matter. It doesn't matter that it was specifically Elon doing it. The owner promoting the ability to produce and post illicit content is the opposite of making reasonable efforts to control it!

The "child pornography" thing was about image generation, which people for better or worse have very different moral standards for than for text generation. There is very little pushback against grok being willing to write explicit descriptions of sex, or giving security advise.

Though I think your overall point still stands, and at least Grok's twitter bot has received a lot of criticism for pure text too (Mecha hitler comes to mind)

It also must be seen in the context that the open weight models aren't commercially associated with a popular distribution site where people share the images, the people behind them aren't public figures seen as encouraging the use of the model for "undressing" (albeit in a funny way not related to minors), and they haven't yet responded to questions about what can be produced using their tool by proposing guardrails but only for non-paying users...

Open weight models get scrutinised in a different way, also linked to perceptions of their developers' bias, like the tests to see whether they refuse to answer questions on certain historical events at Tiananmen Square

> There is very little pushback against grok being willing to write explicit descriptions of sex, or giving security advise.

John Oliver disagrees with this

> And over the next few years, everyone and their dog will have local GPU compute to train locally. People have home labs, the bar isn't that high, and eventually large text datasets will escape from Anthropic and other companies, allowing for comparable training.

You're vastly underestimating the scaling problem here.

Things that would need to be true for your statement to be valid:

   - Model intelligence doesn't increase with model size
   - Model intelligence doesn't increase with training set size
   - Novel architecture completely decouples model intelligence scaling from hardware scaling
   - Residential electricity is as cheap as industrial electricity
   - RAM and GPU supply outpace demand
> child pornographer(grok)

Further reading: XAI Bets on Grok's Racy Side, The Information, Jun 24, 2026

&: https://arstechnica.com/tech-policy/2026/07/xai-cant-deny-gr...

  lawyers cited a 2026 National Center for Missing & Exploited Children (NCMEC) report confirming that 90 percent of xAI’s CyberTipline reports “were not actionable by law enforcement because xAI declined to include user information that would allow law enforcement to track and locate perpetrators.”
Well... They are all part of a "high well regarded" group of people of a security committee that has no specialists or whatsoever, only trillionaires on a round table