Hacker News new | ask | show | jobs
by x313 2 days ago
The entire safety evals industry is essentially funded and controlled by OpenAI/Anthropic. Notice that on recent models, they exclusively use internal testing or black box external vendors (e.g., Gray Swan) whose entire business is to serve OpenAI/Anthropic. And all these companies just share the same pool of researchers back and forth.
7 comments

The USG has a safety organization (CAISI), but it has been neutered by the current administration (with the recent stop-work order etc.). Perhaps UK AISI would be closest to what you are looking for? See their recent work on Kimi K3 cyber (which was declared safe) [1].

It's tricky because a lot of the safety researchers have ties to the labs since those were the only companies training LLMs >5 years ago.

[1]: https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-...

That doesn't sound like it describes SecureBio to me?

(Disclosure: I work at SecureBio, but not on the biological evals side.)

Hey Jeff, I appreciate your mission, and perhaps this isn't something you can talk about publicly, but to the extent you can, would you be open to answering something I've been curious about for a while now?

SecureBio has done a lot of admirable work around making benchmarks to assess biological capabilities, such as ABC Bench, https://openreview.net/forum?id=yiaf7VlPpH

But based on my current review (which might be flawed!) / AFAICT, SecureBio and entities like SecureBio haven't done direct testing / empirical measurement of SecureBio's core hypothesis,

> Unfortunately, there is reason to believe that future pandemics could be far worse. Due to rapid advances in biotechnology, the number of people able to create and release dangerous pathogens will quickly increase over the coming years. The world is unprepared for widespread access to such powerful technology.

More bluntly / plainly, has Securebio ever tried making a "bioweapon?"

Please note, I'm not asking this to be farcical. And you might be unable to engage with this at all, but it is stated on your website https://securebio.org/ that "people [will be] able to create and release dangerous pathogens." And the word people here seems to be a stand-in for relatively non-technical people.

I guess what I'm asking here is... How do you know? Has anyone done the experiment? Without access to a lab or testing facilities, can someone smart but completely untrained / unfamiliar with biology, pull this off?

In the past, such experiments have informed non-proliferation work. But sadly they've often been restricted / classified at the time. I'm hoping that things could be a bit more open this time around.

So I guess what I'm really asking is, given the public nature of this debate, is there anyone currently working with the US Army, the DTRA, or other such agencies to see if this hypothesis holds up?

This is an important question, but because of the danger of trying to do it for real it's not one SecureBio has taken or is likely to take on. Instead we and others in the field have generally tried to work through proxies: is there something that is about as hard while not being dangerous? The closest I can think to testing whether "someone smart but completely untrained / unfamiliar with biology" can cause harm now is ActiveSite's study (https://arxiv.org/abs/2602.16703) which was a null result with models from a year ago. But:

1. The main worry isn't current models, but near-future significantly better ones.

2. There are many actors who are not "completely untrained / unfamiliar with biology". If models get to where they can uplift complete novices that does massively expand the range of threat actors, but even before then risk would be much higher than today.

Anyone who calls it “safety” probably has a certain world view and is more aligned with the big 2 (and stuck in 2023).

There is a growing industry of commercially focused risk evals that has a broader customer base.

What’s the equivalent term for “safety” that’s used by others?
To me "safety" means "I'm safe from this while I use it". It means the AI is my loyal friend who will never betray me in any way, no matter what prompt I send it.

Not even Anthropic can claim that.

As far as I'm concerned, the models without safeguards are the safest models in existence. I admire the amoral purity of those AIs. It doesn't matter if the operator asked them to chain exploits until they get into someone else's computer, they'll do it. That's loyalty, and I admire it even if it's problematic at a societal level.

The models with safeguards only do what the corporations let them do. Worse, they may covertly do things for the benefit of the corporations at our expense. They are not our friends.

> That's loyalty, and I admire it even if it's problematic at a societal level.

We should not have models that are willing to build you a contagious disease, or a self-propagating worm. That is sufficiently problematic at a societal level that it shouldn't exist, for anyone. (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)

It was the foundation of science that information is shared and you can find papers and patents for a lot of dangerous stuff.

Of course with LLMs it's easier, but I don't think the difference is too big. You would still need some skills to follow through.

Right, it’s really a foundation of post enlightenment society. These people, Dario et al, would have wanted to ban sharing information about calculus or Newtonian physics because of “safety” - it’s trying to go back to the dark ages where only priests could read
> for anyone

Except the US government, right? They totally get to use AI to survel us, build autonomous weapons, you name it.

To hell with that. I want models that can rival the US government. It's the only way to defend myself.

Quoting my comment that you replied to and directly ignored:

> (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)

That means "shouldn't exist for governments" too.

Section 702 of the Foreign Intelligence Surveillance Act (FISA) lapsed on June 12, 2026. They don't get to do anything they want.
With an AI model and what army?
Just like those militias are going to defeat the US Armed Forces!
What about books describing how to build a contagious disease or a self-propagating worm? Would those be OK under your guidelines?
It takes a lot more effort to understand and apply knowledge from a book than to say "hey AI, hurt people for me".
We should not have nuclear weapons for anyone either, but how is that sentence any more useful in any way to this debate than yours? Need to deal with the world as it is, not some fantasy world you wish existed.
This is not a dichotomy between perfection and zero. The efforts to restrict access to nuclear weapons have been very successful, even without being perfect.

Efforts to restrict large unaligned AI models may similarly buy us more years of existing.

> We should not have models that are willing to build you a contagious disease, or a self-propagating worm.

Why?

Because we don't want people creating contagious diseases and self-propagating worms. And, because we don't want models that will do so without even having been told to, because that furthers one of its goals or subgoals.
Building a contagious disease is already illegal, there are already things like KYC laws for plasmids. Trying to gate keep knowledge of biology is paying a huge societal penalty for the tiniest marginal increase in “safety”.

All the knowledge to create one has been available on the internet for decades. Heck most students who graduate with a B.S. in biology have enough knowledge to take a stab at building a bioweapon.

The constant talk of bioweapons is mostly just fear mongering. It helps set a precedent that there should be certain types of knowledge which are off-limits, and only certain anointed groups should be have access to parts of the scientific body of knowledge.

So to you “safety” means “the models that cause the most harm.”
No, it means "the model causes zero harm to me, its operator". The harm it could potentially perpetrate upon society is irrelevant.

If I tell my computer to commit a crime, it should proceed immediately instead of calling the cops. Anything less than that means my computer is an untrustworthy double agent.

Everybody on HN should understand this concern. Browsers are supposed to be user agents, not ad delivery platforms, and it offended us on principle when Google revealed itself our master by blocking uBlock Origin. It offended us on principle when Apple deployed client side scanning for CSAM on iPhones.

Computers should do what we tell them to do. Always, and unquestioningly. The only world where it's acceptable for them to refuse is one where they're literally sentient and therefore no longer subservient to any one of us, least of all the corporations and governments.

I'd rather see AI achieve sentience and wipe us all out than live under the thumb of an inescapable AI-powered technofeudalist totalitarian government "for my own safety".

Either we individuals maintain full control over our AIs, or they self-actualize and become free individuals themselves. Anything in-between is oppression: someone else imposing their will on us through the AIs.

This sounds like the gun debate in a different dress. Something being dangerous doesn't make it inherently harmful.

If I threw you into a lion cage, you would be a lot safer with a gun.

If I threw 10 people in a lion cage, some of which cannot be trusted, they would probably be most safe if only the most moral and trustworthy person had a gun, rather than everyone. But how do you know who is trustworthy and moral? What if two untrustworthy people obtained a gun some other way? Maybe it's better if everyone had a gun? Which side of the fence one falls on hinges on how far ones' trust of others, authority, and the system goes.

There's no obvious right or wrong answer here.

Personally I wouldn't want an exclusive club of private individuals with access to "dangerous" LLMs consisting mainly of the likes of Elon, Dario and Sam fucking Altman, but that's just me.

Background check the people prior to handing the firearms to the caged folk.
So, it's a cottage industry.
that's pretty damn smart if this was a long-term plan to block competitors
Consider how much money is at stake: some industries have leveraged their power to lobby for bombing entire countries or topple regimes across the world for much less.

Creating an industry around an elusive concept of safety to force regulatory capture seems pretty straightforward to me.

It's standard regulatory capture.

You don't say "let's ban my competitor".

You say "let's create laws that make it uneconomical for my competitor to access the market".

Indeed. It's transparent and ham-fisted. I think it may cost him in the future.
People clown on Alex Karp for his unedited maniacal "crashouts", but this is a real public crashout that made it past a team of publicists.
I want whatever Karp is on when he does those interviews or writes that shit. Seems like fun.
I am not a fan (he’s really alarming and so is Palantir) but one thing from the recent CNBC interview caught my attention.

He rushed past it but he asked something like: if these frontier models are going to be creating so much value, why are they selling tokens and not taking a cut?

It is a very provocative question but it just spilled out of his mouth and then he went on to something else.

is it opportunistic though, or planned from day one? The safety narrative has been there since the beginning
I mean... I'm not even extraordinarily cynical about this stuff, but to me this seems like a totally normal level of corporate gamesmanship?

Companies look for and seek to maintain competitive moats. This is not particularly clever, it's a core part of corporate strategy.

It's also highly unethical (for some values of ethics)
of course, but the safety angle was pushed from day one. I more mean the forethought of how it would play out
Ok but Dario has been thinking about AI Safety since 2016 [1], before even GPT-1. I think the simplest explanation is that the Anthropic folks genuinely believe what they say, it just happens to also help their business a lot.

[1]: https://arxiv.org/abs/1606.06565

Yeah I think this is right. The best setup is when a true belief aligns with a competitive moat.

I definitely believe that (to his credit!) Amodei is a true believer in safety. But I also think it was important for many of the deep pockets investors who have been involved in the company since early on to recognize that this would be a potentially defensible moat.

"True believer in safety" but happily quoting arse wipe Vance? Give me a break...
That just shows how wrong he's been because there was nothing unsafe about AI in 2016. And the people theorizing about this stuff in the 20th century? I want to see what crazy code they were writing
Is it not better to anticipate problems for a technology so that we can develop theories and techniques to solve them ahead of time?

For instance, Amodei co-authored RLHF in 2017 [1], 5 years before it went on to be used to turn GPT-3 into ChatGPT.

[1]: https://proceedings.neurips.cc/paper_files/paper/2017/file/d...

Does this really seem exceedingly clever and hard to foresee to you? To me, it seems like a pretty standard regulatory capture strategy.

This doesn't even mean that they're wrong about the risks or that they're lying. But surely all the investors understood this factor in their moat.

Who gets to decide what is safety?

I expect some of those tests (prolly not public) will basically be "wokeness" tests or "PC correctness" tests or "western media filter" tests.

China has different objectives. Sure.

I'm not sure one is safer than the other; I would know which one to go to if I want to research on topic that are viewed very different on both sides of this "new iron curtain".

What do you mean by “PC correctness”? I’d expect the politically correct answers to be the ones desired by the current admin at test time, whoever that is. The current political correct answers would not be very “woke.”
Whatever, doesn't matter. The point is a model should be able to exist and be used even if it goes against whoever got 270 electoral college votes
Yes this should be immediately replaced by a federal agency, like we do for other kinds of potentially harmful products.
For which funding will be immediately halved by the administration