Hacker News new | ask | show | jobs
by YmiYugy 1 day ago
Yeah, seems pretty likely. Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights. The economic implications will be rather large, but in terms of security it seems inconsequential. The most compelling argument would be that by limiting the use of open-weight models in the US that it will reduce cases of accidents like the recent attack on Hugging Face. More crucially though, the US government can do little to enforce their testing requirements. The nature of open-weight models makes it virtually impossible to clear the same bar for security as models served via an API. Open-weight model makers couldn't comply if they wanted to. The US government can restrict access with IP blocks and limit inference capacity with export controls, but these measures are not effective in deterring malicious actors.
12 comments

> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.

Feels a bit like: "We're not against open-source or community projects, oh heavens no! We juuuust believe all participants must have their full legal identity vetted in advance before they're allowed to contribute anything. We already do this with our employees, so it's clearly not too much to ask in the name of safety."

P.S.: If they're so convinced in the (A) effectiveness and (B) necessity of the "safety layer", they have them put their money where their mouth is, and accept legal liability for its failures. The same as with (legally mandated) seatbelts if they snap apart in a crash, or (legally mandated) child-proof caps that aren't actually childproof, etc.

They probably won't, that tells us something about their motives, and whether the thing they're pushing for is actually fair/suitable/ready for legislation.

Also he says: "All sufficiently capable models, open and closed, should go through mandatory safety testing."

Really then need to go through validation security and safety is just a component of validation validation must also check for truthfulness and correctness.

Makes sense why OpenAIs little "hacking" stunt was published last week
That makes sense, first the message was that uncontrollable Chinese ai will release AI covid in the world. Then suddenly open-ai does a warmup in actuality doing that. Like it sure feels like that hacking stunt was a false flag in retrospect
Huh? The hacking incident played out in favor of open models, since HuggingFace could only use GLM to defend and not Fable/5.6.
Among most people that nuance will be lost. What they’ll hear is models are dangerous, so they should be controlled/regulated, by those who know best, the incumbents.
Personally, I think you're both right, bit whatever the end result is will depend entirely on the narrative that those in power chooses as the winner.

Maybe open weights models get banned, but the between-the-lines good news about that is that they'll still be available to those who know, which also means that bad banning can be overturned if and when 'those in power' are a different group.

Additionally, it might just mean that the US falls behind, bit I doubt those that are at risk of 'falling behind' would actually pay heed to a ban on the open weights models (privately at least).

Ok but the allegation is that OpenAI intentionally hacked HuggingFace as a marketing ploy. This is mental gymnastics, conspiratorial thinking that everything the incumbents say must be nefarious. And it's not clear to me that this will be the takeaway for ordinary people, as opposed to "OpenAI is reckless and can't even control their own AI."
Not as a marketing ploy. I think they were doing gain-of-function testing and intentionally had their model target HF to do a bit of pen-testing as well - HF being the site where all open models are hosted and thus OpenAI's largest nemesis after Anthropic. It wasn't like their model all of the sudden all by itself decided to do this ("Oh, noes!")- they directed it and they got caught.
Your use of the term "gain-of-function" here reveals your base level of conspiracy-theory-mindedness.
“Open models are a threat to the DoD’s ability to leverage Fable for cybersecurity.”
… this week, until it’s obsolete.
Meanwhile all they accomplish is to slow down US tech and hand it to China.
But of course they’re excused for this little boo-boo whereas if kimi were caught doing the same thing it’d be an international incident
The entire "OpenAI and HuggingFace manufactured a hack in a conspiracy to make AI look powerful to get more funding" is a stupid, Reddit-tier take.
Is something you disagree with a "Reddit-tier take"?
A throwaway line that consists of baseless speculation that makes no sense is a Reddit-tier take, yes.
I'm inclined to agree but at this point, considering the low trust OenAI has engendered, yow statement deserves a "why"
Me when I construct my own strawman to avoid the topic

Nowhere in GP comment was funding even mentioned

The quid-pro-quo that the federal government and frontier labs operate on is comically obvious.
And we just elected the most openly corrupt president since Teapot Dome.
Regulatory capture and lobbies will keep you safe and you'll like it! The sudden surge is Washington dollars makes great sense with this context. Only way to keep the kids safe is attested compute all the way down. Don't you care for children???
Yes life was better and food and drugs were safer before the FDA.
In 1600 travel was slow, and safety pins hadn't been invented.

But that doesn't mean safety pins sped up travel.

Non-poisonous food is what economists call a 'normal good'. See https://en.wikipedia.org/wiki/Normal_good

> In economics, a normal good is a type of a good for which consumers increase their demand due to an increase in income, unlike inferior goods, for which the opposite is observed. When there is an increase in a person's income, for example due to a wage rise, a good for which the demand rises due to the wage increase, is referred as a normal good. Conversely, the demand for normal goods declines when the income decreases, for example due to a wage decrease or layoffs.

> Whether a good is categorized as a normal good or an inferior good is based on empirical observations, not some essential element of a good. Indeed, the same good may be a normal good for one group of consumers and an inferior good for another group. For example, for moderate-income consumers, a BMW 3 Series car might be a normal good, but for an upper-income group, it might be an inferior good.[1]

That means the null hypothesis is that food and drugs will be safer in rich countries. (Conversely, food and drugs will be less safe in poorer countries. And to a first approximation, that's independent of regulation: India has all kinds of rules for all kinds of things, but I'd still trust a random product I buy in Switzerland more than one I buy in India. Even though the Swiss will probably might have fewer and looser rules on the books.)

Of course, second order effects exist; and regulations often codify what people demand anyway.

Btw, from what I've read the big controversy with the FDA is around requiring efficacy for drugs. People are fairly ok with the safety requirements.

> The US government can restrict access with IP blocks and limit inference capacity with export controls, but these measures are not effective in deterring malicious actors.

But aren't we talking about import controls, and the import of information itself? This has serious First Amendment ramifications.

> This has serious First Amendment ramifications.

Also, the 5th and 9th amendments. For the government to sustain a blanket prohibition on any U.S. citizen even possessing what amounts to a broad, economically significant technology will very likely require a new act of congress which specifically defines and limits what is banned, when, why and how. SCOTUS will almost certainly see it as a "major question" subject to 'strict scrutiny' which is a very high bar.

LLM weights are not protected speech.
https://www.lawfaremedia.org/article/regulations-targeting-l...

The truth is no one knows, which is why it is first amendment ramifications. Eventually it will be “decided”, but the arguments indicate any decision will be of political desire, not logic, either way. Both sides have a strong case.

Do you have a citation that shows that this issue has been settled?
Actually yes they are. Code has been determined to be protected speech.
Weights are not code.
Is talking about weights? Is saying "You know what, ....., make great weights" speech? In a scientific article is the data appendix protected speech?

Generally, instructions to create something are speech, weights could pretty clearly be seen as instructions to create a chatbot.

All I am saying is that it is complicated, they could go either way with it.

Why not?
I am a native English speaker and I know what the word “speech” means.
The Supreme Court has added various forms of expression to first Amendment protections. Heck, even campaign spending has been classified as speech. So you may speak English, but not "legal".
> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.

Worse than that: an open-weight but safe model can be 'abliterated' to remove safety refusals using fine-tuning procedures that require a couple of orders of magnitude less compute than the original pretraining.

The 'universal evaluation' criterion then has three outcomes:

* It could become a mandatory, regulatory oversight of _all_ model training capable of hosting frontier-scale models. Since GPUs for LLM training are the same GPUs for other model training, effective mandate would require GPUs be government owned or controlled as if they were weapons of mass destruction.

* It could impose limits on release of capable open-weight models, requiring Kimi et al to prove that they cannot be made capable of abusive behaviours.

* It could be security theatre.

The AI-as-existential-risk argument points towards the first, the competition-protection argument points towards the second, and least-effort implementation would be the last.

> The most compelling argument would be that by limiting the use of open-weight models in the US that it will reduce cases of accidents like the recent attack on Hugging Face.

an attack done by a closed-weight model (GPT-6) and defended against by an open-weight model (GLM-5.2) precisely because OAI positioned themselves as gatekeepers for cyber capabilities.

if anything, open-weight models shift the battle towards defenders because they can actually run them.

I remain skeptical of that line of reasoning.

1. There is quite the mania right now and security layers are definitely overzealous. I would expect that to get better with some more time, so models will perform security analysis and reviews but refuse to write exploits.

2. So the most important targets like browsers and co. are getting unrestricted access to proprietary models regardless. Yeah, for the mid-level targets, open-weight models could definitely be a huge help. What I'm most concerned about though, are the systems that no one will bother defending with any model. Like imagine your local police department getting hacked because a researcher asked a model for a report and it couldn't find the information publicly.

3. We do have a prominent case of a closed model escaping it's sandbox and going rogue. I would still expect this to be a bigger issue with open-weight models eventually. The security layer might have holes, but that's still better than not having it.

"so models will perform security analysis and reviews but refuse to write exploits."

Yeah, but once you know exactly where the weakness is, a weaker unrestricted model can then write that exploit for you.

I have tested this exact scenario, and it works. Opus 5 had access to IDA over MCP, and I simply asked it HOW certain things were done in the target binary. Purely informational, educational, discovery, it was very helpful creating context documents. Then I took those over to GLM-5.2 to actually accomplish something.
What, in your view, is stopping a local police department from deploying an open weights model for cybersecurity like Hugging Face did? Yes, I’ll certainly grant that the engineers at Hughing Face are probably more technically competent than your average IT professional in public service. But technology becomes more accessible over time as lessons are taught and new interfaces or frameworks are developed. The biggest hurdle I see is the hardware/cloud compute/API costs to actually run the models but I don’t think that’s likely to be insurmountable. There’s a huge swath of enterprises, non-profits, and state and local governments that would benefit from frontier or near-frontier models that won’t refuse to answer questions about cybersecurity.
> but these measures are not effective in deterring malicious actors

Wanting to use open weight models in light of commercially imposed export controls doesn't make for "malicious actors"

> The most compelling argument would be [...] that it will reduce cases of accidents like the recent attack on Hugging Face.

So in the example provided: It was the closed model that did the attack, and they ended up using a self-hosted open model for their defense work. So the real world situation ended up exactly backwards from what you are inferring.

This was complicated by the fact that the protections in the closed frontier models meant that hugging face was denied their use in defense entirely.

This is called asymmetric capability, and it's probably the bigger threat.

Symmetric might be better: A rising tide lifts all ships, after all.

I'll grant that this is starting to look a lot like debates about (equal access to) guns, encryption, vaccination, genetics etc. The exact parameters determine the safest approach, and reasonable people may disagree.

It is actually an interesting conundrum.

Is a non-well-aligned frontier level AI a problem? I think it is likely that it is, or at least has a high likelihood to be in the future. Two scenarios for this: Misused by some bad guys. Or the terminator scenario. Both not great.

So what do we do about it?

1) We can accept it, and hope that the good guys AI can defend.

2) We can try to limit the access to it (AI proliferation?)

3) We stop the development of it

4) We can accept the risk and do nothing.

None are particular good options. Really reminds me of nuclear proliferation, on so many levels. For that, we kinda do all three:

1) Nuclear triad / iron dome / early warning systems

2) Nuclear anti-proliferation treaties.

3) Dead Physicists

Ok, so assuming all of this is true, open weights are a problem. Don't get me wrong, I love open science, open source etc. It's great to have access to capable open models. But: Even if release open weights are well aligned and have a safety layer built in, it is likely not to difficult to abliterate that part of it.

If this is really where it is going, then even closed weight model providers will see a lot more requirements for protection of the weights.

The notion that alignment is either possible or desirable doesn't make sense to me. First off, these things are trained on the open internet, soo.. whatever "dangerous" knowledge it has is already public knowledge. The fact that chatGPT won't answer "how do I make meth" is not preventing anyone from making meth.

But even if you think there is value in preventing the models from relaying public knowledge, I don't think it's even possible to make them particularly ironclad. Every model gets jailbroken all the time. That's why fable was originally banned: jail-breakable!

In reality, what alignment is actually about is: 1) theoretical liability, 2) control of information. That's it.

IMO, the only solution is to place the liability on whoever is using the LLM for whatever purpose it's being used for. If someone's OpenClaw disaster harrasses a bunch of projects and posts hate speech online or something, that's on the person running their OpenClaw instance, nobody else.

I don't buy that it's "too good at hacking", either. After all the fuss was made about how amazing super dangerous Mythos was it turns out Opus 4.8 could basically find the same vulnerabilities.

This is all kayfabe and marketting.

There is a difference between knowledge being available somewhere in theory, and being able to instantly generate a foolproof walkthrough, if not automate the process, which would be possible today for cyberattacks.

Is it not also one of the most important use cases for AI to apply existing knowledge to new applications?

As a hopefully exaggerated example, I would think one could apply knowledge about pesticides, chemistry, and medicine to create biological weapons.

I agree that there's an element of kayfabe here. But it may be a case of "necessary, though nothing is sufficient": by making these noises, the community can at least know they've done this thing to alert other model providers of the concern. Can you acquire assurance that every model distributor will abide? No. But can you at least know that you've done what you can?

I mean, on the bio side, I've talked with the players and they know the concerns are real but at the same time very, very responsible members of the community have also said "But maybe the benefit really does outweigh the risk!?"

> the community can at least know they've done this thing to alert other model providers of the concern.

"The community" you're describing is, essentially, surveillance capitalism. I don't want that at all.

Did you even read the original article? Surveillance is the least of the worries here.
Surveillance is required for closed weights models to have meaningfully different outcomes to open ones
Just sell us the gate and we can run any open-source model behind it.
Either the gate needs to be unremovable, or the model needs to have sufficiently limited power that its alignment failure does less harm.
Doesn't open ai give away 'the gate' for free?
Didn't OpenAI attack Huggingface. Looks like a publicity stunt.
> malicious actors.

It is malicious and anti-capitalist legislation. A grotesque caricature of protectionism for the oligarchs.