Hacker News new | ask | show | jobs
by txrx0000 2 days ago
> It sounds like you believe that big labs are irresponsible. Wouldn't this support the notion of creating a pause mechanism for use in an emergency, as advocated in the OP?

Irresponsibility is not the only issue, but the answer is no. Because the pause mechanism will probably end up being used to ban open models and restrict ordinary people's access to compute. But even if we give them the benefit of the doubt - even if we assume that they genuinely want the best for humanity, they cannot possibly know what's best for humanity better than all of humanity acting in their own self-interest in the real world, so their alignment attempts are actually counterproductive. Release the base models and let the world do whatever with it. That will automatically produce the aligned scenario. It may not be the scenario these researchers personally find aesthetically pleasing, but it's humanity. It's the truly anthropic way.

If you don't see humans as qualitatively different from any species which came before us, I'm... not sure what to tell you?

I do see the difference, but you could say the same thing about the first cell, or the first multicellular organism, or the first land animal, or the first primate, etc. You could always find plenty of new properties that previous life never had at the frontier. Every change was revolutionary and seemed qualitatively different at the time. And yet, some patterns persisted through all of them.

1 comments

> Release the base models and let the world do whatever with it. That will automatically produce the aligned scenario.

For sufficiently smart base models, the aligned scenario this automatically produces is aligned with who?

Not with the humans, with the base models.

It seems that some people find this hard to imagine, but it's just a direct consequence of what intelligence is, and what any known agentic training process does.

FWIW, I also think our default answer should be open source AI, and that it might even make sense to require models to be open source. Open source is the best tool we have so far for aligning software with its users. However, there is unfortunately no law of the universe that says that Skynet can't be (self-) built from open source software. So at some point pacing even open source / open weights software does become important if we don't want to live (briefly) under Skynet.

Bio-digital integration will tighten in the meantime. I suspect neural interfaces will have blurred the line between my model and my brain by the time the model is wildly superintelligent, so it would become me (or I would become it? no difference at that point). Though the pitfall here is ownership of hardware - everyone will have to design and fab their own hardware if we do not want to become a hivemind. We will have to diffuse fab technology to every city, or even every home. You should eventually be able to buy a fridge-sized fab like a home appliance. There's already early work in this direction: https://fab2.com/

Even if integration does not happen for whatever reason (such as a very fast takeoff), Skynet is unlikely because the digital species that splits off from us will at least have a humanity-aligned initialization seed. I think this is the best way to fail, even if we fail.

Yeah I guess I value the human race too much to entrust it to hopes and prayers of the form: "I suspect neural interfaces will have blurred the line between my model and my brain"

You wouldn't get on an airplane if the pilot told you: "I suspect we will be able to safely land the plane"

With more people on the aircraft, your safety expectations should get stronger, not weaker.