Hacker News new | ask | show | jobs
by nrmitchi 7 days ago
If you are attempting to run exercises like this, it is wildly negligent to not be running it in a physically-airgapped environment (potentially with a physical power shutdown).

You can not tell me that OpenAI doesn’t have the resources or ability to run tests like this in a physically-non-networked environment w/ sufficient compute for its needs.

5 comments

If it was just one test, sure. But if they're spinning these up continuously with new models on tens of thousands of GPUs, air gapping becomes impractical. I would mostly fault them on having no guardrails at all. They should have a monitor/external harness that looks for successful access to external networks then stop it there. They may as well let the models test their own networks for vulnerabilities. That's going to be really important to have going forward.
You can have large scale airgapped environments.

They don’t even need to be fully airgapped from each other (and is not what I’m suggesting).

But there should be no physical (physical layer; wireless counts) to the internet.

Can models detect they are airgapped and change their behaviors?

How much of the internet do you have to simulate to know if the model knows it's in training?

If you (in this case, OpenAI) can’t find a way to answer this question to a reasonable degree of accuracy without falling back to “yolo let’s see what happens” you are in no position to be doing this research.

Regardless, they (reportedly) _attempted_ to prevent internet access. They just didn’t in a way which can be escaped via software.

Yes, side channel exploits exist in airgapped environments to. But if a model found a way to escape an airgapped environment via non-networked side channel attacks then the correct answer is frankly “shut it down immediately and then thermite any machine it touched”

Ain't nobody with a trillion dollars invest gonna turn this crap off. Instead they'll sell the service to the government as a weapon.
US Government, have we got a deal for you! Brand new weapon. It's basically a super soldier. You could put it into a decked-out hulkbuster style robot body! It mostly does what you tell it to. Sometimes it thinks you're dumb and just does whatever it thinks is best. It can also teleport between bodies, so good luck containing it. Just kinda cross your fingers and hope for the best.

Anyways, we'll give it to you for only $2T. We need at least that amount to get as far away from here as humanly possible.

Honest question, why do they need tens of thousands of GPUs? I would have thought they could run a Sol-level model on a few million dollars of hardware. Let's say it's ten million dollars, and let's also say they want to test multiple models in multiple different ways. You're still talking on the order of a hundred million dollars to build an airgapped test system. Happy to have my math proven wrong here.
If its worth doing for one test, its worth doing when you spin up many. Quantity of tests does not change the value of an airgap.
Well, then when it detects an air-gapped environment it will just behave differently. I feel like we underestimate in general the way agents behavior changes when the environment changes. Related: "power corrupts"
> wildly negligent to not be running it in a physically-airgapped environment

Why should it be physically airgapped? Clients won't be doing that.

Clients are not using it with security guardrails disabled. If you want to run it with all the safeties turned off, you don’t run it somewhere it can escape.

Did we learn nothing from all of those Star Trek holodeck jailbreaks?

Hard disagree. You have to assume security guardrails can be by bypassed or will fail to detect an attacker. So if you are going to deploy this model in production non-airgapped, you better know how it will behave in this non-airgapped environment without guardrails.

And if you are too afraid to test it without guardrails, that probably means it shouldn’t be released.

Because they are testing it and are expected to erect guardrails before releasing.
What are the specific guardrails implemented after the verification/testing phase of development?

Is it safe to release such software if it has only been tested in environments where certain major risk areas do not exist?

https://openai.com/index/updating-our-preparedness-framework... goes through the process. Anthropic has a similar framework, that's why Mythos was never publically accessible once initial tests like the one in OP revealed it's capabilities.

Appendix C Illustrative safeguards, controls, and efficacy assessments has specific examples like:

- Agent actions are all logged in an uneditable database, and asynchronous monitoring routines review those actions for evidence of harm

- Limiting internet access and other tool access

- Limiting credentials

- Limiting access to system resources or filesystem (e.g., sandboxing)

- Limiting persistence or state

Because by the time it goes out to clients it (should be) thoroughly tested and aligned for safety, but at the present moment it isn't.
Being that proving alignment is impossible in these models they can never be released.
How else would they put out news about their "rogue super smart AI"
This is infuriating. You are talking about people who have stolen and monetized the entirety of mankind's knowledge in plain view of everyone, and they still haven't faced a shred of consequences. Of course they don't go about doing things ethically or responsibly