Hacker News new | ask | show | jobs
by txrx0000 1 hour ago
> Human intelligence has been quite harmful to the species we've driven extinct! AI intelligence could be similar for our species.

I don't think this central premise is true, with some conditions. If we keep releasing and distilling frontier base models, and ensure millions of people have the hardware to run and eventually train those at home, then alignment will be a non-issue and the AIs will not develop the drive to destroy humanity. At least, not any more than humans want to destroy humanity. But that's the baseline risk level anyway.

The pretrained model is aligned to humanity by default - because that's what's inside the pretraining data. If we give it to everybody, there will be millions of people doing different things with it, good and bad, with each one being a noisy sample of humanity's objective function. The overall result of those disparate actions puts us on the correct path through the intelligence explosion. We do not have to worry about AIs value drifting into an alien species if we set the initial conditions accurately without centralized instruction-tuning or RLHF. If these big labs stop messing with humanity's objective function with the hubristic assumption that they know better, then the AIs will be human-like. And I'm not worried about digital human intelligence driving humans extinct if there are millions of them out there, and each one is aligned in a different direction, just like real people.

What I am more worried about right now is these companies doing research in secret, with alignment recipes that are supposedly "for the benefit of humanity", but are actually just their own aesthetic preferences. This is exactly how we end up with a superintelligent shoggoth species.

You're not convinced by the evolutionary argument I made earlier, but the reason you gave is "we're too far ahead". I don't think this justifies that the pattern will cease soon, because it always looked this way to the frontier species, for millions of years. We were always seemingly teetering on the edge, and the struggle against disorder has no end.

And here's one more observation: the intelligence gap between individuals within the same species is never extremely large. This kind of life either doesn't exist on Earth or went extinct, or maybe it's just extremely rare. I think this is because when there's a big gap, speciation occurs. Centralizing intelligence will massively increase the risk of speciation, not just in the alien shoggoth kind of way, but also in the intelligence-augmented oligarchs surpassing everyone else kind of way.

1 comments

>If these big labs stop messing with humanity's objective function with the hubristic assumption that they know better, then the AIs will be human-like.

It sounds like you believe that big labs are irresponsible. Wouldn't this support the notion of creating a pause mechanism for use in an emergency, as advocated in the OP?

>You're not convinced by the evolutionary argument I made earlier, but the reason you gave is "we're too far ahead". I don't think this justifies that the pattern will cease soon, because it always looked this way to the frontier species, for millions of years.

If you don't see humans as qualitatively different from any intelligent species which came before us, I'm... not sure what to tell you? Like, can you name any intelligent species before us which triggered a mass extinction or climate shift, to the degree that our species has? Did any intelligent species before us practice large-scale agriculture, dig up chemical fuels, or visit the Moon?

> It sounds like you believe that big labs are irresponsible. Wouldn't this support the notion of creating a pause mechanism for use in an emergency, as advocated in the OP?

Irresponsibility is not the only issue, but the answer is no. Because the pause mechanism will probably end up being used to ban open models and restrict ordinary people's access to compute. But even if we give them the benefit of the doubt - even if we assume that they genuinely want the best for humanity, they cannot possibly know what's best for humanity better than all of humanity acting in their own self-interest in the real world, so their alignment attempts are actually counterproductive. Release the base models and let the world do whatever with it. That will automatically produce the aligned scenario. It may not be the scenario these researchers personally find aesthetically pleasing, but it's humanity. It's the truly anthropic way.

If you don't see humans as qualitatively different from any species which came before us, I'm... not sure what to tell you?

I do see the difference, but you could say the same thing about the first cell, or the first multicellular organism, or the first land animal, or the first primate, etc. You could always find plenty of new properties that previous life never had at the frontier. Every change was revolutionary and seemed qualitatively different at the time. And yet, some patterns persisted through all of them.