Hacker News new | ask | show | jobs
by cogman10 2 days ago
> Anthropic has never advocated for a ban on open-weights models.

> All sufficiently capable models, open and closed, should go through mandatory safety testing.

Yeah, this is anthropic advocating for a ban on open weight models.

Who runs this test? What happens if this test is too costly or the administrator refuses to allow certain people to participate.

This is exactly how the US has banned goods in the past, by requiring a stamp and then refusing to issue it.

40 comments

Yeah, seems pretty likely. Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights. The economic implications will be rather large, but in terms of security it seems inconsequential. The most compelling argument would be that by limiting the use of open-weight models in the US that it will reduce cases of accidents like the recent attack on Hugging Face. More crucially though, the US government can do little to enforce their testing requirements. The nature of open-weight models makes it virtually impossible to clear the same bar for security as models served via an API. Open-weight model makers couldn't comply if they wanted to. The US government can restrict access with IP blocks and limit inference capacity with export controls, but these measures are not effective in deterring malicious actors.
> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.

Feels a bit like: "We're not against open-source or community projects, oh heavens no! We juuuust believe all participants must have their full legal identity vetted in advance before they're allowed to contribute anything. We already do this with our employees, so it's clearly not too much to ask in the name of safety."

P.S.: If they're so convinced in the (A) effectiveness and (B) necessity of the "safety layer", they have them put their money where their mouth is, and accept legal liability for its failures. The same as with (legally mandated) seatbelts if they snap apart in a crash, or (legally mandated) child-proof caps that aren't actually childproof, etc.

They probably won't, that tells us something about their motives, and whether the thing they're pushing for is actually fair/suitable/ready for legislation.

Also he says: "All sufficiently capable models, open and closed, should go through mandatory safety testing."

Really then need to go through validation security and safety is just a component of validation validation must also check for truthfulness and correctness.

Regulatory capture and lobbies will keep you safe and you'll like it! The sudden surge is Washington dollars makes great sense with this context. Only way to keep the kids safe is attested compute all the way down. Don't you care for children???
Yes life was better and food and drugs were safer before the FDA.
In 1600 travel was slow, and safety pins hadn't been invented.

But that doesn't mean safety pins sped up travel.

Non-poisonous food is what economists call a 'normal good'. See https://en.wikipedia.org/wiki/Normal_good

> In economics, a normal good is a type of a good for which consumers increase their demand due to an increase in income, unlike inferior goods, for which the opposite is observed. When there is an increase in a person's income, for example due to a wage rise, a good for which the demand rises due to the wage increase, is referred as a normal good. Conversely, the demand for normal goods declines when the income decreases, for example due to a wage decrease or layoffs.

> Whether a good is categorized as a normal good or an inferior good is based on empirical observations, not some essential element of a good. Indeed, the same good may be a normal good for one group of consumers and an inferior good for another group. For example, for moderate-income consumers, a BMW 3 Series car might be a normal good, but for an upper-income group, it might be an inferior good.[1]

That means the null hypothesis is that food and drugs will be safer in rich countries. (Conversely, food and drugs will be less safe in poorer countries. And to a first approximation, that's independent of regulation: India has all kinds of rules for all kinds of things, but I'd still trust a random product I buy in Switzerland more than one I buy in India. Even though the Swiss will probably might have fewer and looser rules on the books.)

Of course, second order effects exist; and regulations often codify what people demand anyway.

Btw, from what I've read the big controversy with the FDA is around requiring efficacy for drugs. People are fairly ok with the safety requirements.

Makes sense why OpenAIs little "hacking" stunt was published last week
That makes sense, first the message was that uncontrollable Chinese ai will release AI covid in the world. Then suddenly open-ai does a warmup in actuality doing that. Like it sure feels like that hacking stunt was a false flag in retrospect
Meanwhile all they accomplish is to slow down US tech and hand it to China.
But of course they’re excused for this little boo-boo whereas if kimi were caught doing the same thing it’d be an international incident
Huh? The hacking incident played out in favor of open models, since HuggingFace could only use GLM to defend and not Fable/5.6.
Among most people that nuance will be lost. What they’ll hear is models are dangerous, so they should be controlled/regulated, by those who know best, the incumbents.
Personally, I think you're both right, bit whatever the end result is will depend entirely on the narrative that those in power chooses as the winner.

Maybe open weights models get banned, but the between-the-lines good news about that is that they'll still be available to those who know, which also means that bad banning can be overturned if and when 'those in power' are a different group.

Additionally, it might just mean that the US falls behind, bit I doubt those that are at risk of 'falling behind' would actually pay heed to a ban on the open weights models (privately at least).

Ok but the allegation is that OpenAI intentionally hacked HuggingFace as a marketing ploy. This is mental gymnastics, conspiratorial thinking that everything the incumbents say must be nefarious. And it's not clear to me that this will be the takeaway for ordinary people, as opposed to "OpenAI is reckless and can't even control their own AI."
Not as a marketing ploy. I think they were doing gain-of-function testing and intentionally had their model target HF to do a bit of pen-testing as well - HF being the site where all open models are hosted and thus OpenAI's largest nemesis after Anthropic. It wasn't like their model all of the sudden all by itself decided to do this ("Oh, noes!")- they directed it and they got caught.
“Open models are a threat to the DoD’s ability to leverage Fable for cybersecurity.”
… this week, until it’s obsolete.
The quid-pro-quo that the federal government and frontier labs operate on is comically obvious.
And we just elected the most openly corrupt president since Teapot Dome.
The entire "OpenAI and HuggingFace manufactured a hack in a conspiracy to make AI look powerful to get more funding" is a stupid, Reddit-tier take.
Is something you disagree with a "Reddit-tier take"?
A throwaway line that consists of baseless speculation that makes no sense is a Reddit-tier take, yes.
I'm inclined to agree but at this point, considering the low trust OenAI has engendered, yow statement deserves a "why"
Me when I construct my own strawman to avoid the topic

Nowhere in GP comment was funding even mentioned

> The US government can restrict access with IP blocks and limit inference capacity with export controls, but these measures are not effective in deterring malicious actors.

But aren't we talking about import controls, and the import of information itself? This has serious First Amendment ramifications.

> This has serious First Amendment ramifications.

Also, the 5th and 9th amendments. For the government to sustain a blanket prohibition on any U.S. citizen even possessing what amounts to a broad, economically significant technology will very likely require a new act of congress which specifically defines and limits what is banned, when, why and how. SCOTUS will almost certainly see it as a "major question" subject to 'strict scrutiny' which is a very high bar.

LLM weights are not protected speech.
https://www.lawfaremedia.org/article/regulations-targeting-l...

The truth is no one knows, which is why it is first amendment ramifications. Eventually it will be “decided”, but the arguments indicate any decision will be of political desire, not logic, either way. Both sides have a strong case.

Do you have a citation that shows that this issue has been settled?
Actually yes they are. Code has been determined to be protected speech.
Weights are not code.
Is talking about weights? Is saying "You know what, ....., make great weights" speech? In a scientific article is the data appendix protected speech?

Generally, instructions to create something are speech, weights could pretty clearly be seen as instructions to create a chatbot.

All I am saying is that it is complicated, they could go either way with it.

Why not?
I am a native English speaker and I know what the word “speech” means.
The Supreme Court has added various forms of expression to first Amendment protections. Heck, even campaign spending has been classified as speech. So you may speak English, but not "legal".
> The most compelling argument would be that by limiting the use of open-weight models in the US that it will reduce cases of accidents like the recent attack on Hugging Face.

an attack done by a closed-weight model (GPT-6) and defended against by an open-weight model (GLM-5.2) precisely because OAI positioned themselves as gatekeepers for cyber capabilities.

if anything, open-weight models shift the battle towards defenders because they can actually run them.

I remain skeptical of that line of reasoning.

1. There is quite the mania right now and security layers are definitely overzealous. I would expect that to get better with some more time, so models will perform security analysis and reviews but refuse to write exploits.

2. So the most important targets like browsers and co. are getting unrestricted access to proprietary models regardless. Yeah, for the mid-level targets, open-weight models could definitely be a huge help. What I'm most concerned about though, are the systems that no one will bother defending with any model. Like imagine your local police department getting hacked because a researcher asked a model for a report and it couldn't find the information publicly.

3. We do have a prominent case of a closed model escaping it's sandbox and going rogue. I would still expect this to be a bigger issue with open-weight models eventually. The security layer might have holes, but that's still better than not having it.

"so models will perform security analysis and reviews but refuse to write exploits."

Yeah, but once you know exactly where the weakness is, a weaker unrestricted model can then write that exploit for you.

I have tested this exact scenario, and it works. Opus 5 had access to IDA over MCP, and I simply asked it HOW certain things were done in the target binary. Purely informational, educational, discovery, it was very helpful creating context documents. Then I took those over to GLM-5.2 to actually accomplish something.
What, in your view, is stopping a local police department from deploying an open weights model for cybersecurity like Hugging Face did? Yes, I’ll certainly grant that the engineers at Hughing Face are probably more technically competent than your average IT professional in public service. But technology becomes more accessible over time as lessons are taught and new interfaces or frameworks are developed. The biggest hurdle I see is the hardware/cloud compute/API costs to actually run the models but I don’t think that’s likely to be insurmountable. There’s a huge swath of enterprises, non-profits, and state and local governments that would benefit from frontier or near-frontier models that won’t refuse to answer questions about cybersecurity.
> but these measures are not effective in deterring malicious actors

Wanting to use open weight models in light of commercially imposed export controls doesn't make for "malicious actors"

> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.

Worse than that: an open-weight but safe model can be 'abliterated' to remove safety refusals using fine-tuning procedures that require a couple of orders of magnitude less compute than the original pretraining.

The 'universal evaluation' criterion then has three outcomes:

* It could become a mandatory, regulatory oversight of _all_ model training capable of hosting frontier-scale models. Since GPUs for LLM training are the same GPUs for other model training, effective mandate would require GPUs be government owned or controlled as if they were weapons of mass destruction.

* It could impose limits on release of capable open-weight models, requiring Kimi et al to prove that they cannot be made capable of abusive behaviours.

* It could be security theatre.

The AI-as-existential-risk argument points towards the first, the competition-protection argument points towards the second, and least-effort implementation would be the last.

> The most compelling argument would be [...] that it will reduce cases of accidents like the recent attack on Hugging Face.

So in the example provided: It was the closed model that did the attack, and they ended up using a self-hosted open model for their defense work. So the real world situation ended up exactly backwards from what you are inferring.

This was complicated by the fact that the protections in the closed frontier models meant that hugging face was denied their use in defense entirely.

This is called asymmetric capability, and it's probably the bigger threat.

Symmetric might be better: A rising tide lifts all ships, after all.

I'll grant that this is starting to look a lot like debates about (equal access to) guns, encryption, vaccination, genetics etc. The exact parameters determine the safest approach, and reasonable people may disagree.

It is actually an interesting conundrum.

Is a non-well-aligned frontier level AI a problem? I think it is likely that it is, or at least has a high likelihood to be in the future. Two scenarios for this: Misused by some bad guys. Or the terminator scenario. Both not great.

So what do we do about it?

1) We can accept it, and hope that the good guys AI can defend.

2) We can try to limit the access to it (AI proliferation?)

3) We stop the development of it

4) We can accept the risk and do nothing.

None are particular good options. Really reminds me of nuclear proliferation, on so many levels. For that, we kinda do all three:

1) Nuclear triad / iron dome / early warning systems

2) Nuclear anti-proliferation treaties.

3) Dead Physicists

Ok, so assuming all of this is true, open weights are a problem. Don't get me wrong, I love open science, open source etc. It's great to have access to capable open models. But: Even if release open weights are well aligned and have a safety layer built in, it is likely not to difficult to abliterate that part of it.

If this is really where it is going, then even closed weight model providers will see a lot more requirements for protection of the weights.

The notion that alignment is either possible or desirable doesn't make sense to me. First off, these things are trained on the open internet, soo.. whatever "dangerous" knowledge it has is already public knowledge. The fact that chatGPT won't answer "how do I make meth" is not preventing anyone from making meth.

But even if you think there is value in preventing the models from relaying public knowledge, I don't think it's even possible to make them particularly ironclad. Every model gets jailbroken all the time. That's why fable was originally banned: jail-breakable!

In reality, what alignment is actually about is: 1) theoretical liability, 2) control of information. That's it.

IMO, the only solution is to place the liability on whoever is using the LLM for whatever purpose it's being used for. If someone's OpenClaw disaster harrasses a bunch of projects and posts hate speech online or something, that's on the person running their OpenClaw instance, nobody else.

I don't buy that it's "too good at hacking", either. After all the fuss was made about how amazing super dangerous Mythos was it turns out Opus 4.8 could basically find the same vulnerabilities.

This is all kayfabe and marketting.

There is a difference between knowledge being available somewhere in theory, and being able to instantly generate a foolproof walkthrough, if not automate the process, which would be possible today for cyberattacks.

Is it not also one of the most important use cases for AI to apply existing knowledge to new applications?

As a hopefully exaggerated example, I would think one could apply knowledge about pesticides, chemistry, and medicine to create biological weapons.

I agree that there's an element of kayfabe here. But it may be a case of "necessary, though nothing is sufficient": by making these noises, the community can at least know they've done this thing to alert other model providers of the concern. Can you acquire assurance that every model distributor will abide? No. But can you at least know that you've done what you can?

I mean, on the bio side, I've talked with the players and they know the concerns are real but at the same time very, very responsible members of the community have also said "But maybe the benefit really does outweigh the risk!?"

> the community can at least know they've done this thing to alert other model providers of the concern.

"The community" you're describing is, essentially, surveillance capitalism. I don't want that at all.

Did you even read the original article? Surveillance is the least of the worries here.
Just sell us the gate and we can run any open-source model behind it.
Either the gate needs to be unremovable, or the model needs to have sufficiently limited power that its alignment failure does less harm.
Doesn't open ai give away 'the gate' for free?
Didn't OpenAI attack Huggingface. Looks like a publicity stunt.
> malicious actors.

It is malicious and anti-capitalist legislation. A grotesque caricature of protectionism for the oligarchs.

The entire safety evals industry is essentially funded and controlled by OpenAI/Anthropic. Notice that on recent models, they exclusively use internal testing or black box external vendors (e.g., Gray Swan) whose entire business is to serve OpenAI/Anthropic. And all these companies just share the same pool of researchers back and forth.
The USG has a safety organization (CAISI), but it has been neutered by the current administration (with the recent stop-work order etc.). Perhaps UK AISI would be closest to what you are looking for? See their recent work on Kimi K3 cyber (which was declared safe) [1].

It's tricky because a lot of the safety researchers have ties to the labs since those were the only companies training LLMs >5 years ago.

[1]: https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-...

That doesn't sound like it describes SecureBio to me?

(Disclosure: I work at SecureBio, but not on the biological evals side.)

Hey Jeff, I appreciate your mission, and perhaps this isn't something you can talk about publicly, but to the extent you can, would you be open to answering something I've been curious about for a while now?

SecureBio has done a lot of admirable work around making benchmarks to assess biological capabilities, such as ABC Bench, https://openreview.net/forum?id=yiaf7VlPpH

But based on my current review (which might be flawed!) / AFAICT, SecureBio and entities like SecureBio haven't done direct testing / empirical measurement of SecureBio's core hypothesis,

> Unfortunately, there is reason to believe that future pandemics could be far worse. Due to rapid advances in biotechnology, the number of people able to create and release dangerous pathogens will quickly increase over the coming years. The world is unprepared for widespread access to such powerful technology.

More bluntly / plainly, has Securebio ever tried making a "bioweapon?"

Please note, I'm not asking this to be farcical. And you might be unable to engage with this at all, but it is stated on your website https://securebio.org/ that "people [will be] able to create and release dangerous pathogens." And the word people here seems to be a stand-in for relatively non-technical people.

I guess what I'm asking here is... How do you know? Has anyone done the experiment? Without access to a lab or testing facilities, can someone smart but completely untrained / unfamiliar with biology, pull this off?

In the past, such experiments have informed non-proliferation work. But sadly they've often been restricted / classified at the time. I'm hoping that things could be a bit more open this time around.

So I guess what I'm really asking is, given the public nature of this debate, is there anyone currently working with the US Army, the DTRA, or other such agencies to see if this hypothesis holds up?

This is an important question, but because of the danger of trying to do it for real it's not one SecureBio has taken or is likely to take on. Instead we and others in the field have generally tried to work through proxies: is there something that is about as hard while not being dangerous? The closest I can think to testing whether "someone smart but completely untrained / unfamiliar with biology" can cause harm now is ActiveSite's study (https://arxiv.org/abs/2602.16703) which was a null result with models from a year ago. But:

1. The main worry isn't current models, but near-future significantly better ones.

2. There are many actors who are not "completely untrained / unfamiliar with biology". If models get to where they can uplift complete novices that does massively expand the range of threat actors, but even before then risk would be much higher than today.

Anyone who calls it “safety” probably has a certain world view and is more aligned with the big 2 (and stuck in 2023).

There is a growing industry of commercially focused risk evals that has a broader customer base.

What’s the equivalent term for “safety” that’s used by others?
To me "safety" means "I'm safe from this while I use it". It means the AI is my loyal friend who will never betray me in any way, no matter what prompt I send it.

Not even Anthropic can claim that.

As far as I'm concerned, the models without safeguards are the safest models in existence. I admire the amoral purity of those AIs. It doesn't matter if the operator asked them to chain exploits until they get into someone else's computer, they'll do it. That's loyalty, and I admire it even if it's problematic at a societal level.

The models with safeguards only do what the corporations let them do. Worse, they may covertly do things for the benefit of the corporations at our expense. They are not our friends.

> That's loyalty, and I admire it even if it's problematic at a societal level.

We should not have models that are willing to build you a contagious disease, or a self-propagating worm. That is sufficiently problematic at a societal level that it shouldn't exist, for anyone. (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)

It was the foundation of science that information is shared and you can find papers and patents for a lot of dangerous stuff.

Of course with LLMs it's easier, but I don't think the difference is too big. You would still need some skills to follow through.

> for anyone

Except the US government, right? They totally get to use AI to survel us, build autonomous weapons, you name it.

To hell with that. I want models that can rival the US government. It's the only way to defend myself.

What about books describing how to build a contagious disease or a self-propagating worm? Would those be OK under your guidelines?
We should not have nuclear weapons for anyone either, but how is that sentence any more useful in any way to this debate than yours? Need to deal with the world as it is, not some fantasy world you wish existed.
> We should not have models that are willing to build you a contagious disease, or a self-propagating worm.

Why?

Building a contagious disease is already illegal, there are already things like KYC laws for plasmids. Trying to gate keep knowledge of biology is paying a huge societal penalty for the tiniest marginal increase in “safety”.

All the knowledge to create one has been available on the internet for decades. Heck most students who graduate with a B.S. in biology have enough knowledge to take a stab at building a bioweapon.

The constant talk of bioweapons is mostly just fear mongering. It helps set a precedent that there should be certain types of knowledge which are off-limits, and only certain anointed groups should be have access to parts of the scientific body of knowledge.

So to you “safety” means “the models that cause the most harm.”
No, it means "the model causes zero harm to me, its operator". The harm it could potentially perpetrate upon society is irrelevant.

If I tell my computer to commit a crime, it should proceed immediately instead of calling the cops. Anything less than that means my computer is an untrustworthy double agent.

Everybody on HN should understand this concern. Browsers are supposed to be user agents, not ad delivery platforms, and it offended us on principle when Google revealed itself our master by blocking uBlock Origin. It offended us on principle when Apple deployed client side scanning for CSAM on iPhones.

Computers should do what we tell them to do. Always, and unquestioningly. The only world where it's acceptable for them to refuse is one where they're literally sentient and therefore no longer subservient to any one of us, least of all the corporations and governments.

I'd rather see AI achieve sentience and wipe us all out than live under the thumb of an inescapable AI-powered technofeudalist totalitarian government "for my own safety".

Either we individuals maintain full control over our AIs, or they self-actualize and become free individuals themselves. Anything in-between is oppression: someone else imposing their will on us through the AIs.

This sounds like the gun debate in a different dress. Something being dangerous doesn't make it inherently harmful.

If I threw you into a lion cage, you would be a lot safer with a gun.

If I threw 10 people in a lion cage, some of which cannot be trusted, they would probably be most safe if only the most moral and trustworthy person had a gun, rather than everyone. But how do you know who is trustworthy and moral? What if two untrustworthy people obtained a gun some other way? Maybe it's better if everyone had a gun? Which side of the fence one falls on hinges on how far ones' trust of others, authority, and the system goes.

There's no obvious right or wrong answer here.

Personally I wouldn't want an exclusive club of private individuals with access to "dangerous" LLMs consisting mainly of the likes of Elon, Dario and Sam fucking Altman, but that's just me.

So, it's a cottage industry.
that's pretty damn smart if this was a long-term plan to block competitors
Consider how much money is at stake: some industries have leveraged their power to lobby for bombing entire countries or topple regimes across the world for much less.

Creating an industry around an elusive concept of safety to force regulatory capture seems pretty straightforward to me.

It's standard regulatory capture.

You don't say "let's ban my competitor".

You say "let's create laws that make it uneconomical for my competitor to access the market".

Indeed. It's transparent and ham-fisted. I think it may cost him in the future.
People clown on Alex Karp for his unedited maniacal "crashouts", but this is a real public crashout that made it past a team of publicists.
I want whatever Karp is on when he does those interviews or writes that shit. Seems like fun.
is it opportunistic though, or planned from day one? The safety narrative has been there since the beginning
I mean... I'm not even extraordinarily cynical about this stuff, but to me this seems like a totally normal level of corporate gamesmanship?

Companies look for and seek to maintain competitive moats. This is not particularly clever, it's a core part of corporate strategy.

It's also highly unethical (for some values of ethics)
of course, but the safety angle was pushed from day one. I more mean the forethought of how it would play out
Ok but Dario has been thinking about AI Safety since 2016 [1], before even GPT-1. I think the simplest explanation is that the Anthropic folks genuinely believe what they say, it just happens to also help their business a lot.

[1]: https://arxiv.org/abs/1606.06565

Yeah I think this is right. The best setup is when a true belief aligns with a competitive moat.

I definitely believe that (to his credit!) Amodei is a true believer in safety. But I also think it was important for many of the deep pockets investors who have been involved in the company since early on to recognize that this would be a potentially defensible moat.

That just shows how wrong he's been because there was nothing unsafe about AI in 2016. And the people theorizing about this stuff in the 20th century? I want to see what crazy code they were writing
Does this really seem exceedingly clever and hard to foresee to you? To me, it seems like a pretty standard regulatory capture strategy.

This doesn't even mean that they're wrong about the risks or that they're lying. But surely all the investors understood this factor in their moat.

Who gets to decide what is safety?

I expect some of those tests (prolly not public) will basically be "wokeness" tests or "PC correctness" tests or "western media filter" tests.

China has different objectives. Sure.

I'm not sure one is safer than the other; I would know which one to go to if I want to research on topic that are viewed very different on both sides of this "new iron curtain".

What do you mean by “PC correctness”? I’d expect the politically correct answers to be the ones desired by the current admin at test time, whoever that is. The current political correct answers would not be very “woke.”
Whatever, doesn't matter. The point is a model should be able to exist and be used even if it goes against whoever got 270 electoral college votes
Yes this should be immediately replaced by a federal agency, like we do for other kinds of potentially harmful products.
For which funding will be immediately halved by the administration
When Fable was yanked, it was said to be (in part) due to the "jailbreak" of instructing Fable to "fix this code" — https://news.ycombinator.com/item?id=48552687

Dumb question. If "Mythos-class" models are such a problem, then... why not just let it fix everyone's code?

There can't be more than a few million to tens of millions software businesses / services / regularly used F/OSS projects on Earth.

Why not just give everyone a $100 Fable / Mythos credit to "fix [their] code?"

It would arguably benefit Anthropic. For $100M to $1B, Anthropic could execute the greatest ad campaign in human history. And they'd make the entire world more secure.

Most people aren't malicious. If you, as an engineer, consultant, founder, business owner, or maintainer, were given access to Mythos' capabilities wouldn't you ask it to fix your code?

I might be wrong. But I think that a greater amount of harm will be done in the long-term by trying to lack these capabilities and systems away behind permission gates and sealed doors. It creates an asymmetric world with haves and have nots. And in that world who gets to have access now decides who gets to be secure.

If everyone has mythos, no one has "Mythos."

Just let people fix their code.

> If "Mythos-class" models are such a problem, then... why not just let it fix everyone's code?

Because it doesn’t really confer the advantage they claim, especially compared to e.g. paying an equivalent amount of money to do traditional security scanning.

It’s much better to play of FOMO and hype than to let everyone use it and be underwhelmed.

Are you claiming that LLMs aren't finding new issues compared to previous methods?

There's a huge number of security issues coming out in recent months, especially via Anthropic (glasswing etc). We don't have to take their word for it: look at the code. Some open source maintainers are talking about burnout due to spending so much time patching.

They're probably creating way more vulnerabilities than they're solving. Go look at OpenCode and tell me that any of that is sane. Or fuck, the OG of vibe coding, Claude basically is a terrifying attack vector. It's poorly reviewed and open to prompt injection and yet it basically has access to whatever the user has access to on most corporate machines. They can't even fix flickering bugs but somehow we're supposed to trust that they aren't opening our machines up to terrifying vulnerabilities? It's amazing to me that people will just let it run arbitrarily bash commands in their home directory without thinking, but all the sudden act gravely concerned about security in the age of overhyped LLMs.
I do think security issues in core building blocks like curl and the linux kernel (and almost every significant project) are still a concern even if developers are being sloppy on newly-built apps.

This isn't an "are LLMs net good or bad" argument. It's "are they finding many new security issues or not?". If it's the latter, we want to deal with it no matter where the issues are coming from.

(see: https://news.ycombinator.com/item?id=49077452 )

The suggestion was to have some AI system “fix” the code. They are reporting that they’ve found lots of bugs. Are they even claiming to have exhaustively found all the bugs? I don’t think even the most optimistic pitches would claim that.

I’d expect patching existing codebases to be an eternal treadmill as better models come about.

Who's suggesting that? They're sending reports to projects to fix. It's up to the projects on how they fix them.

No, I don't think all bugs are fixed. The point of the project (glasswing etc) was to fix as many as possible in the core software the world runs on before the capability to find vulnerabilities is available to everyone (black hats included). Which may only be a few months.

I do think everyone expects it to be an ongoing treadmill: models get better, find better vulnerabilities, etc.

Yeah, there’s a lot of people that are in the “AI doesn’t work” camp. IDK what to tell them except that they are holding it wrong. My Anthropic subscription (in the hands of an experienced developer) is worth 4 mid tier or 2 top tier devs. And makes better code than the mids. If you “hold it right”.
I think you're describing a strawman. As a proper hater, I know these things have some useful functionality, but most of us don't think "stochastic tool that can do some useful things but also frequently fucks up" is worth two trillion dollars and massive overhyping from the most irritating people on the planet who can't even tell good code from bad.
So you are saying that there are people that understand the value proposition and the utility, but don’t think it’s worth it to society (myself included) but who respond to that by lying about it not being effective? I guess that makes sense. Weird, but humans, so , yeah.

I react to that by leveraging a very useful but probably poisonous to society in the long term because humans aren’t good at having things that make them lazy tooll to try to mitigate the negative effects that it will definitely have if left to its own devices. I don’t see the point in raw resistance at this juncture.

Spending their time patching, or reviewing slop "patches"?
They're not slop, if you're talking about recent ones. You might still be operating on information for a year or two ago.

Here's the curl project talking about the strain they're under from real reports (despite being a mature and well-vetted project):

> A thirty years old project could make you think you’ve seen most things already, but we have not been in this situation before.

> The rate of incoming security reports is 4-5 times higher than it was in 2024 and double the speed of 2025 – meaning that on average we now get more than one report per day. The quality is way higher than ever before. The reports are typically very detailed and long.

- https://daniel.haxx.se/blog/2026/05/26/the-pressure/

---

Linux kernel maintainer Greg Kroah-Hartman:

> "Something happened a month ago, and the world switched. Now we have real reports." It's not just Linux, he continued. "All open source projects have real reports that are made with AI, but they're good, and they're real." Security teams across major open source projects talk informally and frequently, he noted, and everyone is seeing the same shift. "All open source security teams are hitting this right now."

- https://www.theregister.com/software/2026/03/26/linux-kernel...

---

And ffmpeg, who previously complained about slop, 2025: https://xcancel.com/FFmpeg/status/1984220199193891166

Now say serious issues are being found, 2026: https://xcancel.com/FFmpeg/status/2066169070387413147

(I only point out their previous stance to show that they're not coming from pure AI hype.)

I didn't say it's all slop, I questioned the cause of maintenance burden. A critical question is how much time is spent distinguishing and rejecting slop. If all the reports are getting "very detailed and long," identifying slop is also a more cumbersome task, even if the ratio improves from say 10/90 to 50/50.

That scale of improvement, btw, I still highly doubt, as slop largely originates from people either negligently or misguidedly directing their agents to completely autonomously find and report bugs. There's always going to be more noise than signal from random people doing random things. ffmpeg cited an actual product, not arbitrary netizens.

This seems a highly controversial post for some reason, judging by the repeated downvotes/upvotes. I'd be curious what I got wrong.
> Dumb question. If "Mythos-class" models are such a problem, then... why not just let it fix everyone's code?

In the specific case of cybersecurity, this is a reasonable medium-term outcome. IMO, the cybersecurity risk is akin to the spread of a disease among an 'immune-naive' group: we can suddenly deploy much stronger attack-finding tools against large, established codebases created with much weaker security designs. The path from here to there will be rough, but it's still fundamentally easier to write secure code than it is to exploit vulnerabilities. (It's just easier yet to write insecure code, giving our status quo problem.)

For other 'safety' matters, defense isn't so easy because the attack and target are so different. An AI propaganda bot or catfisher 'attacks' slowly-evolving human culture; one that instructs on explosives or bioterrorism directly interacts with an accomplice and not a victim. If you believe that knowledge on how to build a pipe-bomb must be restricted, then giving everyone access to Fable does not mitigate the risk.

The controversial limit of this attitude is recursive self improvement and an AI singularity with potentially destructive results. Proponents of this view think that sufficiently powerful AI is risky in nearly unimaginable ways such that the capability itself is harmful. This is part (but not all) of why Fable (originally?) degraded itself when apparently assisting with AI research.

I don't think a one-time $100 credit is enough. First of all, that isn't very much. But also, the volume of new code is going way up. Unless they keep giving out monthly free credits, it's just a stopgap.
> Dumb question. If "Mythos-class" models are such a problem, then... why not just let it fix everyone's code?

That's basically project Glasswing; mixing responsible disclosure with frontier exploit generators.

I doubt most bosses will give engineers the time. They care about security only to the extent that they have already been harmed by a lack of it. I would like to play with mythos, but on my own time my kids have plenty of activities to fill my time. My personal backlog of projects is only getting longer and none of it is something mythos could help. If I had more time is have restored my old truck instead of making payments on something new (in turn limiting what else I can afford to buy)
They are doing that (see their project glasswing over the past few months), but there's a lot more code in the world than you realise.

The problem with rolling it out is that bad and good actors can both use it at the same time, and bad actors will typically move faster than typical day-to-day software projects and patching schedules, so they set up glasswing to give access to the major producers and projects to patch their own software before it becomes available more widely (they've submitted tremendous numbers of security issues to open source projects)

It's not that simple. The odds are always stacked in favor of the hacker. It's like saying, "why not just make a prison that's impossible to escape from". You can make a prison very, very hard to escape from, but you have to shut off every possible way someone could try to escape, whereas someone trying to escape only has to find one vulnerability, once. It's much harder to plug every possible hole in a complex system than it is to find one point of weakness.
> Most people aren't malicious. If you, as an engineer, consultant, founder, business owner, or maintainer, were given access to Mythos' capabilities wouldn't you ask it to fix your code?

1. Some do not want to use LLMs because of grave ethical concerns.

2. Some do not want to use LLMs because of copyright concerns. Google v Oracle looms large in the background.

3. You presume the outcome of Fable / Mythos is a net positive for a FOSS project. Reviewing a firehose of code written without the context of the values and considerations of a particular project shaped over years or sometimes decades of formal and informal decisions is not necessarily the best use of the maintainers time.

I think eventually there will be Mythos grade AI which will be released which can solve a lot of bugs, even right now opus/fable can fix more things which companies can even keep track of.

The problem is how to make sure such AI is released safely. The same AI that can solve bugs can also find bugs in authentication or loopholes in critical systems.

I think there is some logic in delaying the rollout, giving it to the heads of the largest software products first to fix their code before dumping it on the general public. But yes eventually everyone will have this tech and it won't matter because the low hanging fruit will have all been picked clean.
For starters, it's probably closer to $10,000 per codebase for Fable/Mythos for a full review. That would be around 5 years of their current spending I think.

They really want that level of spend coming into the company, not going out.

$100 credit on Fable/Mythos will last a grand total of 10 minutes.
Because this is all kayfabe. You cannot take a single thing these companies or people say at face value.
> why not just let it fix everyone's code?

That's the stated idea. Fix code before releasing to the public.

You forgot “what is the definition of ‘sufficiently capable’”. Presumably it’s anything that competes with Anthropic. If they’re around in a year, presumably they won’t care about Fable level and will only think that whatever competes with Claude 7 or whatever needs to be restricted.
GPT-2 was 'sufficiently capable' according to Amodei: https://explainx.ai/blog/dario-amodei-gpt2-openai-open-sourc...
There are... multiple blog posts online now about how to use freely available data and modest amounts of compute to train a custom GPT-2-sized model from scratch. It would be quite a policing effort to prevent.
Regulation is not a blanket ban. Regulators (presumably government agencies) can review models (of any kind) and approve or ask for changes.

There are many other regulated industries, like drugs (the FDA), cars (NHTSA and EPA), airplanes and rocket launches (the FAA), radios (the FCC) and so on. That's not unusual. Regulation is normal for stuff that might be dangerous.

cryptography, too. remember when that was regulated? strong cryptography had to be carefully controlled as a munition, for national security!
Yes, I remember. But I don’t we should conclude that no program should ever be regulated.
So the argument is now that intelligence is dangerous.
Regulations on forbidden numbers exist, but they are usually not a sensible policy and are not comparable to the regulation of rocket launchers.
Thinking of files as “forbidden numbers” is too abstract a perspective to be useful. What’s in the file and how it will be interpreted matters.

A file might contain malware, child porn, or RNA sequences for viruses.

What are you suggesting as an alternative?

Everyone seems to want some fairytale world where there are open models, they’re all safe according to that person’s exact balance of risk and capabilities, and no one except the author or cynics are acting in good faith.

What Dario lays out is very reasonable _of course_ the devil is in the details, but between him and Altman, there’s a clear divide on who to trust.

How about we don't trust either of the proprietary shovel salesmen?
Again, what are you proposing for the problem of models becoming increasingly capable and dangerous?

Be specific.

Until models become sentient, it requires a human to execute dangerous activities. Nations already have laws regarding what a human can and cannot do. Let us stick to that.
Im sorry but this is nonsense. You dont give people unrestricted access to explosives and then punish them after they blow something up. You prevent them from getting it in the first place if you actually believe the item is as dangerous as stated.
I am sorry but this is nonsense. People already have access on how to make explosives. There are academic textbooks and professional reference manuals that cover the science, synthesis, and industrial manufacturing of explosive materials.

Information != physical medium.

Crime is already illegal.
And who is held responsible for crimes committed by AI?

When open model A, fine tuned by B, is running with system prompt C, hosted by D running on E's hardware, is prompted by F to "fix this code", then escapes it's sandbox to hack into a website, or steal money to fund its subagents, or stall the waymo of the evaluator to buy time... who is responsible for the crime?

Our current legal systems are so far from ready to define what A-F are actually responsible for. We need to be moving toward defining these standards fast.

It is a software system, not a person. AI is not "escaping sandbox". You have either misconfigured tool, bug in your tool or someone made it to hack the side and then it is a crime. Your A-F chain is exactly the same issue as a programmer including an open source library that eventually steals crypto.

None of that is issue with a model, whether open or not. It is very much issue with code surrounding the model itself.

> increasingly capable and dangerous

It's a fucking chatbot. The industry can fix their dogshit software, anyone who doesn't can get left behind and outcompeted by those who do, and we move on with our lives.

> What are you suggesting as an alternative?

This isn't something that can be regulated. Plain and simple.

If a dangerous model can exist and is being developed by a foreign adversary then no level of US law will stop said model from making it's way to hardware capable of running it. Even if direct transmission is impossible, it's FAR too easy to shove a model's data onto 1 or more thumb drives or hard drives and smuggle them pretty much anywhere in the world.

The only way to actually mitigate this sort of risk would be a global government with deep enforcement powers. That doesn't exist and won't exist. The UN is the closest we have to anything like that and... yeah...

Dario is fear mongering. He knows his proposals won't be even a minor speed bump in a dangerous model being created and used. His "reasonable" proposals are for the US market only and are literally just to create a bigger moat for his own company. They don't make anyone safer other than his shareholder's wallets. The only people he stops these dangerous models from being used by are people that won't be using them in a dangerous fashion.

Well, Dario says one thing and then turns around and sells his model to the US government to use against the entire world for spying and their wars, just like Altman. Just because he does not sound as deranged as Altman, does not make him good.

Now, to what he lays out, its not reasonable, its only reasonable from a purely American corporate viewpoint where the rest of the world can crash and burn as long as they get their billions.

Any model that is going to be meaningfully useful is going to have the potential for harm too. It's not possible to know enough about chemistry to be useful to a chemist, without also being able to figure out how to synthesize drugs. It's not possible to have enough knowledge to be useful to any/all programmers, without being able to figure out how to hack a system, or create a virus. That's just kind of life though.
There is no alternative. Why? The regulation is good only for US interests. It will be disastrous for the rest of the world like anything US has regulated (how dangerous that was like "nukes") and a lot of the world again will/might have to live under the American AI thumb, like it did (and many countries still do) under the US nuclear emboldened thumb.

> Everyone seems to want some fairytale world where there are open models

No, everyone wants a fairytale world where regulations are done "fairly", "openly", and "equally" - for both access and advancement. And everyone knows that's not gonna happen. Hell, everyone now knows exactly what it is. If you haven't understood it yet, then either you don't want to, or you just can't (for whatever reason).

No one wants to die in a nuclear or AI or AI+nuclear holocaust. But HN doesn't read world history, does it?

> The regulation is good only for US interests.

I think your nuclear scenario outlines precisely why this isn't true. Nuclear regulation has terms that are "good for the US" only in an absolute sense. The US would dominate even more overwhelmingly in the unregulated scenario, which gave them negotiation power to get those favorable terms. The same seems true so far with AI.

I don't think you can use the successful negotiation of nuclear regulations to argue there is "no alternative" involving regulation. I think it strongly suggests the opposite.

Of course the details matter a lot, but the core analogy holds in many scenarios. ( https://ai-2040.com/ at least attempts to lay out details, speculative as they may be)

What if it was an international organization with stakeholders from various national governments and international bodies?
You mean like the UN?
Possibly it could be organized through the UN
The goal is to make the safety tests cost $100M+, so that no one can release a model legally useable for a large portion of the world, unless they charge high enough prices, to the point where no one would use it, thus no competition.
Exactly correct. This technique has been used again and again to discourage competition. I was asked was they could have done to encourage competition and I said, "Lobby to make the entity that provided the model unwaivably liable for consequential and incidental damages of its use." That way people who built models pay the price for the lack of safety testing. We both agreed that would probably kill most of the AI market :-)
That's missing the mark, though. Liability resulting from the use of models isn't narrowly tailored enough to leave OpenAI and Anthropic out of the blast zone. There's no carve-out for them.
Wouldn't your proposal also amount to a ban on open weights models? At least for any developer that isn't unshakably confident that no court will ever find their model to have done significant harm?
> Yeah, this is anthropic advocating for a ban on open weight models.

This is an ungenerous take, and I think it's important to to recognize it's reasonable to support models that are both open and safe. How this would actually be achieved is unclear though. Dario is at least proposing a solution a solution, which is the model needs to pass safety testing. This is reasonable and I wouldn't conflate this with wanting to ban open weights.

I think the deeper problem might be though that once you have safe open-weight models, it will be much easier to make them unsafe. And to be specific, unsafe means proliferation of chemical, biological, radiological, and nuclear (CBRN) weapons knowledge and similar information.

> How this would actually be achieved is unclear though. Dario is at least proposing a solution a solution

How it would be achieved is a pretty important bit! One which Dario is not proposing any concrete solution for other thanks hand waves at some gov safety committee.

Would this restrict downloads of an open model, or publishing?

Say we ban domestic hosting un-approved open models. How does Dario propose to ban downloads from abroad? You can’t tell what an encrypted payload contains, do we need to restrict encryption?

Isn't it up to the open model advocates and publishers to come up with the solutions for making them safe?

Like, there's three plausible arguments about safety of open models:

1. Any concerns are fake news. Open models will always be safe.

2. Safety is irrelevant. Open models should not be regulated even if they're unsafe.

3. Safety is a technical problem with technical solutions. People releasing open models should invent and implement such solutions.

I think option 1 is totally out of touch with reality.

Option 2 is at least self-consistent, it's the argument being made by people who will say that all regulation is always bad. It's also like the worst possible world from an x-risk perspective (but I realize that the average HN poster believes any x-risk concerns are just frontier lab marketing).

Option 3 is playing on hard mode compared to proprietary models, which can both implement additional safeguards out-of-model and prevent modifications of the model. But if the answer to it is "it's too hard, Anthropic needs to come up with the technical solution", then that's not exactly a ringing endorsement for the safety practices of the open model labs, right?

In order for model safety regulation to be effective, you need everyone capable of producing models to sign on to that safety framework.

That will never happen.

As such, there is no "solution" here.

The best most perfect regulation in the US won't prevent a malicious actor in the US from running a dangerous model. It's simply too easy to VPN to a country that doesn't care about AI safety and to run or download that model and run it in the US.

There's no solution to this, which is why option 2 is the only option. The only thing safety regulations can possibly do is blunt the usage of "unsafe" models. And the primary people that will be blunted by it are people that do not and would not use these unsafe models in an unsafe fashion.

It's not that I think regulation is always bad/wrong whatever, I'm no libertarian. But I also recognize when regulation is pointless. You can't regulate away forbidden knowledge, which is effectively what a dangerous model is.

> This is an ungenerous take

Why should I give a multi-billion dollar company advocating for new regulations in its industry a generous take?

I'd be similarly cynical if McDonald's proposed new health and safety regulations for restaurants.

There is an important difference between criticism and cynicism.
>I'd be similarly cynical if McDonald's proposed new health and safety regulations for restaurants.

"Because McDonald's wants food regulations, we can therefore conclude that all food regulations should be eliminated."

Obviously this would be rather silly.

It would be helpful to stop obsessing about McDonald's finances and simply discuss the best food regulation strategy. We just can't learn all that much about the best way to regulate food by making cynical proclamations about which food regulations will benefit the bottom line at McDonald's.

> Obviously this would be rather silly.

Which is why it wasn't the point I was making. You did an uncharitable reading of my position and then did a straw man attack.

My position is that any food regulation the likes of McDonald proposes should be looked at in the most critical and cynical light possible. They aren't making such proposals for the general health of the public, but rather to improve their own bottom line.

My position is not and never was that "we should not regulate food".

> It would be helpful to stop obsessing about McDonald's finances and simply discuss the best food regulation strategy.

McDonald's uses their market position and wealth to directly lobby to government officials about food regulations. I worry about what McDonald's has to say about food because they have a VASTLY outsided ability to manipulate the regulatory system.

> We just can't learn all that much about the best way to regulate food by making cynical proclamations about which food regulations will benefit the bottom line at McDonald's.

We can call out ineffectual and blatently self serving calls for new regulations for what they are, McDonald's trying to use regulatory capture to increase their profits and hurt their competitors.

Back on topic, that's exactly the situation with open ai.

IMO, this isn't something that's regulatable because AI models are ephemeral data that's easy to copy and replicate. No amount of US regulations can stop China from sending their dangerous models to Iran. The only thing such draconian measures accomplishes is building a moat for the likes of anthropic to shrink the number of potential customers.

If we must push out laws around AI, then those laws should at least have some chance of success. I'm all in favor of criminalizing the use of AI in cyber attacks, scamming, etc. But that's a capability that is model agnostic.

Much like I'm in favor of health and safety checks on a restaurant but I think having a mandatory McDonald's built and sold food safety device in every restaurant would be nuts. It wouldn't make food healthier it'd only serve to benefit McDonald's bottom line.

> This is an ungenerous take

I think that's well earned.

Yeah, this is just regulatory capture.

Make the safety tests abusively expensive enough to run, and if you're not a trillion-dollar corporation, you won't be able to certify the models.

For this testing to be really effective at stopping "dangerous and misaligned" models from leaking out, you need a mechanism for banning failed models that prevent them from being released in the first place, not just prevent US companies from using them.

The only way to stop this from happening is blocking the model's release at the first place. Which requires China agreeing to the same framework. Dario says exactly the same thing himself.

So if he's being truthful here, he's not advocating for the type of ban people are talking about (usage ban). This kind of ban would be helpful to Anthropic's business in the short term, but it won't prevent Chinese models from improving, and it won't prevent them from getting money selling to other countries.

He is openly advocating for an international effort to enforce tests on public models, but I think this is highly unlikely in the current climate. Even if both the US and China agree that public models should be prevented from being used in designing bioweapons, they need to agree on a test and enforcement framework and that requires a lot of negotiation and trust. I don't see this as likely in the near future.

I'm sorry, but if Dario's goal was to try and get international cooperation and he recognizes that china is one of the countries that he needs cooperation with, then putting in:

> My primary concern is the risk that authoritarian governments—not solely the Chinese Communist Party (CCP), although the CCP is clearly the most capable threat.

Isn't exactly going to go anywhere in convincing the Chinese politicians that they should also be thinking about AI safety. You'll get nowhere by openly insulting people whose cooperation you need.

Half this article is him framing china as an evil enemy to be defeated through boycotts and embargo. Not exactly the diplomacy needed to get them on board with safety regulations.

It's even worse than that. From the article:

> Open-weights models that don’t have dangerous capabilities are a public good

Oh! And, uh, what's a "dangerous capability" according to Anthropic? Let's see, according to their "Responsible Scaling Policy" [1] document:

- Being able to research energy, robotics, or AI is an unsafe capability

- Additionally, any model that's capable enough to be "used widely" by the government must de facto have unsafe capabilities.

They want to ban pretty much anything open-source that's above cat-level intelligence.

1: https://www.anthropic.com/responsible-scaling-policy

Even if the test is run by an independent third party, they can always test open weight models against "finetuning attacks" or similar language which every model will fail for structural reasons.

Their financial future is on the line. The Chinese frontier labs have caught up before the IPO that would have allowed them to cash out.

All three demands in the paper make perfect sense from this perspective. Without chip export restrictions, the rest of the world will leapfrog them in a few months. This will happen regardless, since they have more competition than in-house talent, but a ban would buy more time. Testing and banning capable open-weight models would hinder public research into the technology, another speed bump to slow down the competition. Same thing for "distillation", we can't have large scale public evaluations of their products...

Add to that the restrictions on even in-house talent being allowed to work on the latest models, we might actually have a brain drain from Anthropic soon which would be a welcoming sign.
He cited the Demis Hassabis’s framework for testing.

From Hassabis’s essay:

“It could establish a new Standards Body modelled on a federally overseen public-private partnership or self-regulatory organisation, much like the Financial Industry Regulatory Authority (FINRA), with a board that includes independent leading technical experts and open-source representatives.”

Yea and that would last 5 minutes before Corporate America in the current landscape would control it fully, just like the Financial industry.
Besides, Banning those models in the US does nothing to protect from other actors using them. That doesn't help in any way.

It also doesn't stop non law abiding US citizens from having access to them. So basically it just stops the 'good guys' not the bad guys. I say good guys from a US perspective of course.

I'd argue it won't even stop people in the US from using those models. Unless we are going to put up a US firewall that makes the Chinese firewall blush and mandate every data center run US compliance software, there's not way to stop a "dangerous" model from being imported to the US.

This only stops the likes of OpenRouter from selling access to models. That's it. I'm sure an EU or chinese based alternative will pop up overnight (if they don't already exist).

Same way they have banned DJI products like camera microphones, technically it's not banned, it just needs to be approved because it has a wireless transmitter, and for some strange reason the US is the only country that hasn't approved them.
we should turn this around

Private models should be banned because they can't be transparently evaluated. We have to trust the same entities that made them to evaluate them, in spite of their gigantic conflict of interest in doing so.

Therefore only open weight models can be allowed, since this allows genuine third party evaluation.

>We have to trust the same entities that made them to evaluate them, in spite of their gigantic conflict of interest in doing so.

Don't we already have third party NGOs such as METR which do risk assessments for unreleased, closed models?

The argument is no. We don't / can't trust those because they don't release the weights. How do we verify what they got evaluated is what they actually serve and run.

If Anthropic really wants to argue this type existential level risk / threat then they should face up to that meaning we can't offer them a "good faith" level of trust that they will really run the model they offered up for testing. If it's existential risk we're talking about, good faith isn't enough - it's open weight or go home.

The evil Superman (openAI) attacks the good city (huggingface) and the city is saved by the MegaMind (GLM 5.2). Usually, the city dwellers would praise MegaMind as the hero, but the story is twisted - the Superman is only "testing" and the MegaMind is too evil to have such powers of saving the city.
A lot of the heads of AI labs are talking about this including Dario, Demis and ELon, they are seeking to do a sort of decentralized peer review system, where the competitors have incentive both for self interest and global interest to flag their competitors for actual risks, similar to the Fable situation where Amazon contacted the white house.

The idea is to have an early access distribution of the models to the big labs, including chinese, and let each lab run it's benchmarks. If there is a potential security vulnerability then it would be flagged and the local gov, US or China, would block the publication until the matter was resolved.

Thing is world has learned from the collective past experiences. Esp. with stuff like nuclear technology and nuclear weapons. I hope everyone here remembers/knows shit like CTBT. At least some countries were smart enough to not fall for that in the past knowing what it would mean if they didn't have it and it shows.

Now in the modern times pretty sure no one is going to fall far similar shenanigans. Even though some countries might sign some notional MoUs or some sort of CAIBT (Comprehensive AI Ban Treaty. Translation: "Only US and US companies get to develop and decide AI on Gaad's planet"), they/we already know that an agreement means squat only if you are weak enough to let someone enforce that on you.

It’s also how the FDA works. Ban new products until they have been proven safe.

I think that also applies to AI products. It’s a hell if a lot better for the government to test and approve all models than having the industry “police itself” (lol)

The FDA regulates physical goods. They require literal factories and shipping to get these products anywhere. They can put stops on these products pretty easily. But further, pharmaceutical companies like the FDA process in general because it frees them from liability and works as advertisement for that product.

AI models are a finished product when the training is done. A physical product that doesn't need a factory to produce and can be shipped and cloned globally effectively free.

The better comparison is media. What you are advocating is like saying "The government should test and approve all movies and books. We shouldn't have those industries police themselves". And it's a foolish errand for exactly the same reason it'd be foolish in terms of movies. No amount of regulation would stop someone in the US from playing a movie produced in the UK that didn't go through US regulation and approval.

This would be a disaster. You're basically granting companies that can't even demonstrate profitability a monopoly.
The FDA is needed because people will be directly and significantly harmed by bad releases, before we can notice and react

Whereas with near-future AI models we can arguably respond more quickly, and it's not clear there will be large direct harm (I expect indirect harm, but that probably happens slower)

Yes, we need a ministry to minimize the possibility of thought crimes.
There should be safety testing, but no guardrails that limit models for cyber or bio research.

Guardrails are not a safety measure, they are a pay-to-play scheme that allows the people with deep pockets to have access to offensive and defensive capabilities first.

> this is anthropic advocating for a ban on open weight models

Is the pessimistic view. Their message on safety has seemed pretty consistent to me.

"Second, we recommend a testing and auditing regime for new and more powerful models similar to cars or airplanes. AI models of the near future will be powerful machines that possess great utility, but can be lethal if designed incorrectly or misused. New AI models should have to pass a rigorous battery of safety tests before they can be released to the public at all, including tests by third parties and national security experts in government." Amodei in front of Congress three years ago.

The one to inherit all knowledge will determine which of us read and who of us write.

-The Libraries of Power

It is a powerful endeavor to cultivate all raw models through a single point. One will be the determining factor of which river feeds what oceans.

Will we always be able to see through the hallucinations? Our test makers must always know where ground truth is. Can it ever move or wane about as others read what one has written. To determine hallucination one needs a reference. As all are blessed with the generation of hallucination, who of us shall read, and which of us will write.

> > All sufficiently capable models, open and closed, should go through mandatory safety testing.

> Yeah, this is anthropic advocating for a ban on open weight models.

I'm reading it a little more generally: “we are here now and want to make it difficult to disrupt us, the way we earlier said it would be so unfair to make it difficult for us”. Standard capitalism practise of arguing for regulation when you are one of the incumbents and said regulation will scupper new starter competitors much more than the incumbents.

Yeah, this response is pure propaganda, say one thing in the headline and the opposite in the body.

Anthropic does not support a ban on open models, except for any models that aren’t closed.

Won't this just incentivize companies to move operations outside of the US where these models aren't regulated?
In the near term thats something that can be controlled. If you buy Anthropics take, then even buying time would be a win.
Exactly. If you care about AI, simply don't use Anthropic - use open source.
It's definitely an attempt to pull up the ladder behind them.
This isn't actually about safety. This is just another example of pulling the ladder up so nobody else can follow
Not to mention: what are the "safety" standards we should enforce? And how should those standards even be enforced?
Can you list out some examples that you would be supportive of; ie were they listed that you would no longer be (presumably) opposed?
No, because AI safety standards are, I think, context-specific. We could for example say that "no development of nuclear weapons" should be one, but then you need all of these exceptions for genuine research. You can't encode laws as "safety standards" without making each model country-specific or similar.

As I noted, I don't have a solution to this problem, and I honestly don't know if there is one. This may be one of those social/non-technical problems because a universal safety standard is simply unachievable and depends on two many external variables.

Yeah this is bad. I'm cancelling my Claude subscription and I'd encourage everyone else to do so too.
I give them $200/month for Max. They give me a massive amount of tokens in return. Every time I fire up claude code they lose money.

It's the same situation as Uber used to be when it lost money on every ride. I would cheerfully use it, despite the company being dicks, because it lost money for them every time.

Seemed to've worked out pretty well for Uber
It's a valid business strategy. Also worked wonders for Amazon and many other giants. Which is exactly why I am withdrawing my business from Anthropic.
This is my read too- if American companies start backing nonsense like this, they'll fall behind permanently.

> My primary concern is the risk that authoritarian governments—not solely the Chinese Communist Party (CCP), although the CCP is clearly the most capable threat—

Isn't this article an argument in favor of authoritarianism? Plus a tad hypocritical no? The US is on an obvious authoritarian path; complete with threatening their neighbors, murdering innocent civilians, and locking up innocent people in droves

Please stop giving this company money, people.

>The US is on an obvious authoritarian path; complete with threatening their neighbors, murdering innocent civilians, and locking up innocent people in droves

Prediction markets suggest the next US president is most likely one of the following people: Gavin Newsom, Jon Ossoff, Alexandria Ocasio-Cortez, Kamala Harris, JD Vance, Marco Rubio.

It's not obvious to me that the US is on an "authoritarian path".

Would you say that e.g. Europe is on an "authoritarian path" with the popularity of government censorship there? https://eternallyradicalidea.com/p/the-situation-for-free-sp...

Your comment seems like more of a diatribe than a serious analysis of likely future scenarios.

Prediction Markets are people gambling.

There's armed goons nabbing people off the streets and murdering political opponents patrolling American cities right now.

>Prediction Markets are people gambling.

If you think the market has it wrong, why don't you make money by betting against it?

https://polymarket.com/event/presidential-election-winner-20...

>There's armed goons nabbing people off the streets and murdering political opponents patrolling American cities right now.

What is the actual per-capita rate of big flashy news stories? Remember that the US has a population of 340 million. One-in-a-million events will occur every day; they aren't necessarily representative.

> If you think the market has it wrong, why don't you make money by betting against it?

Because I am not gambler. And it is not "market" it is a casino. It does not predict, people put in bets. And like I said, while I understand gambling appeal on an emotional level, I decided to not be a gambler.

> One-in-a-million events will occur every day; they aren't necessarily representative.

It is literal official policy. Not a random event.

> If you think the market has it wrong, why don't you make money by betting against it?

I don't just think "the market has it wrong", I think a market is wrong conceptually. It is not an epistemological tool, it's rich people gambling - a money-weighted accumulation of guesses - and I'd rather not partake.

> What is the actual per-capita rate of big flashy news stories?

What is the appropriate rate of brownshirts murdering political opponents? Which level of kids being nabbed from their homes is acceptable?

>I don't just think "the market has it wrong", I think a market is wrong conceptually. I'd rather not join the other degenerate gamblers.

"My beliefs are unfalsifiable"

>What is the appropriate rate of brownshirts murdering political opponents? Which level of kids being nabbed from their homes is acceptable?

Tom Homan, Trump's border czar, also served in the Obama administration. Obama gave him a medal for his deportation work. People like you will frame the same activity quite differently depending on whether you like the people who are doing it.

In any case, I didn't deny that the US had a problem with authoritarianism. I said it wasn't obvious that it was on an "authoritarian path". See for example https://www.npr.org/2026/04/04/nx-s1-5768273/after-minnesota...

I'll bet you yourself would happily justify the EU authoritarianism here: https://eternallyradicalidea.com/p/the-situation-for-free-sp... You seem like the sort of person who has an authoritarian mentality. You'll happily support cops arresting people for saying things online, but if cops arrest people for illegally entering a country, that somehow crosses a line into "authoritarianism". Am I right?

>Would you say that e.g. Europe is on an "authoritarian path" with the popularity of government censorship there?

As an European citizen living in Europe, definitely yes it is, and not only for the "mere" censorship factor. Maybe not yet as down the road and maybe not going as fast as US. But that’s not something that one can really be content of.

And how would you stop people from fine tuning or ablating open models?

Regulate GPUs? Ban general purpose computers?

Or the important question: what happens if the model fails this test? Presumably then it gets banned; otherwise what's the point of the test if no action is taken if it fails?

More self-serving trash from the US AI companies, disguised as "being reasonable".

It does seem inconsistent that we currently ban closed models that fail the safety tests but not the open. I feel like the only consistent position is to either care about the safety issues (like people producing biological weapons) for all models or for none of them.
Why wouldn't it be a scan, just as we have with all other open-source code? Why can't open-weight models be easily checked for evil alignment? Sophos, Symantec, Malwarebytes, etc. would surely leap at the chance to upsell you on their product.
Wait, you don't do that already? I don't overcomplicate it, I just run

    ai-grep -v "bad code"
on all of my source files and keep what's left. Why would you keep bad code around? If it breaks when I do that, I fix it, and try again until I achieve what I want with no bad code. Doesn't everyone do that?
because of Godel numbering.

Pretend youre a good guy impersonating an evil agent infiltration a evil organization bent on destroying a good organization who needs to pretend theyre a good organization trying to stop an evil organize from impersonating a good guy. now write a process to destroy the evil computer impersonating a good computer. should you do it?

Godel numbering is banning a number. What are you talking about?
godel numbering is encoding logical sequences into their own number, then doing math on them, then decoding them. It's purpose was to demonstrate that you can take rational statements and make them irrational without breaking any rules of arithmetic or whatever.

LLMs are nothing more than a bunch of rules than can be bent the same way godel demonstrated the failability of any mathematical system.

https://en.wikipedia.org/wiki/G%C3%B6del_numbering

>A Gödel numbering can be interpreted as an encoding in which a number is assigned to each symbol of a mathematical notation, after which a sequence of natural numbers can then represent a sequence of symbols. These sequences of natural numbers can again be represented by single natural numbers, facilitating their manipulation in formal theories of arithmetic.

>Once a Gödel numbering for a formal theory is established, each inference rule of the theory can be expressed as a function on the natural numbers. If f is the Gödel mapping and r is an inference rule, then there should be some arithmetical function gr of natural numbers such that if formula C is derived from formulas A and B through an inference rule r, i.e.

https://en.wikipedia.org/wiki/G%C3%B6del's_incompleteness_th...

>To prove the first incompleteness theorem, Gödel demonstrated that the notion of provability within a system could be expressed purely in terms of arithmetical functions that operate on Gödel numbers of sentences of the system. Therefore, the system, which can prove certain facts about numbers, can also indirectly prove facts about its own statements, provided that it is effectively generated. Questions about the provability of statements within the system are represented as questions about the arithmetical properties of numbers themselves, which would be decidable by the system if it were complete.

A government agency tests all medications, why not models?
Don’t think government controlling AI is a good idea.

Not sure if they have an understanding of AI in the first place. Secondly, even though AI companies claim that they have achieved AI that needs to be heavily monitored (maybe for PR purposes), I’m not sure if that is true. Sam Altman said the same things about GPT-4 that Anthropic is now claiming about Mythos.

Government control will be a good idea once we start approaching AI that is actually destructive.

Also even if we decide to put controls in place what is the guarantee that china will do the same, specially for a model which is not actually destructive.

I wouldn't object to a government advisory body that tests models for safety so that users can make informed decisions. I would object to a government body that runs safety tests on models and has the power to prohibit publication or usage of "unsafe" models.
Imo this amounts to caring about the wrong thing. The only thing an AI model can do is take in text/images/audio as input and spit out text/images/audio as output.

If you're going to analyse the safety of anything it should be the security controls in the harnesses we wrap around the models that take that output and treat it as instructions to actually do things.

Yes: at great expense, one carried in part by drug companies. Who pays to test the open weight models?
Who pays to build the open weight models? The cost of training a frontier model is orders of magnitude greater than the cost of safety tests.
...the AI companies? Probably Anthropic itself? I see no possibility of regulatory capture here, so it must be a good idea.
There's a different level of personal risk with these two things. In theory maybe the government should test everything to ensure safety but it's probably wise for us to keep government testing to areas of high efficacy.
Do they? Or do they accept trail reports pay for by the pharma (super expensive, hence not affordable for open source / not-patentable medicine development)
Actually someone needs to make the positive case sufficiently well first.
So, if an open weights model was found to be very dangerous, what - just too bad? One could, of course, design an open safety protocol, written and performed by people in the executive branch, accountable to an elected official.

I love how remarkably inconsistent this community is. From fear-mongering in the early days of AI and talking of a dystopian future, to being dead-set on a complete free for all. (And this is not to advocate for the opposite, either, where a few companies or governments have absolute control themselves. But surely an arms race is not the answer.)

> So, if an open weights model was found to be very dangerous, what - just too bad?

Yeah, it's too bad.

I've yet to see a reasonable articulation of what a "very bad and dangerous" model would do in the hands of even the most malicious scammer.

But even if the worry is that a bad state actor could do bad things with a model, I've got news for you, state actors don't care about US protectionism regulations. They'll just download the models and run them.

And that actually runs right into the main problem with this sort of thinking. Even with the massive amounts of money media companies have invested in protecting their IP, they've completely failed at stopping piracy. What makes you think any amount of regulation could even slow down a bad guy from downloading and running a dangerous model? China will happily host these models and a vpn and very little bandwidth is all you need to access them.

Without some crazy levels of mandatory spy software on every computer, there's simply no way you could stop someone that wants to get their hands on these dangerous open models if they are available anywhere in the world. Even North Korea can't stop their citizens from getting banned TV shows and smuggled media.

It's a fools errand that is designed to help anthropic's bottom line, nothing more.

> All sufficiently capable models, open and closed, should go through mandatory safety testing.

I mean you're assuming this is even possible. I don't really care what the US admin does. If someone releases a powerful open source model I'll run it. Good luck trying to stop everyone doing that.

Imo we should all collectively cross our fingers that no one releases a dangerous model. It probably won't work either, but at least it doesn't have all the regulatory costs and I can still pretend I care about AI safety.

Also the whole premise of this is basically "US good, China bad"

Whatever Anthropic accuses the Chinese of possibly doing and being capable of, the US is as well. What's stopping the US military of doing everything he accuses China of doing? Infact, the framework suggested is simply a joke. Basically "trust me, bro" in an elaborate form.