| >Nukes are weapons of mass destruction that can only destroy. Intelligence is dual-use. Perhaps a better analogy, instead of "nukes", would be "nuclear technology". Just like intelligence, nuclear technology is dual-use: Capable of both energy generation which fuels prosperity, and also mass destruction. Just like intelligence, nuclear technology demands a healthy level of caution and vigilance. >If we look at all organisms on Earth, species that are more intelligent are also more prosperous. This is true from bacteria, to insects, to animals, to humanity. Intelligence itself is not harmful. Human intelligence has been quite harmful to the species we've driven extinct! AI intelligence could be similar for our species. >There will be new adversarial games when everyone's smarter (bio/cyber offense/defense), but those games are always symmetrical in the long run We've had multiple near misses with nuclear technology, such as the Cuban Missile Crisis. I think you're presuming an inevitability which just doesn't exist. Same way a person living in the year 1500 would have no hope of accurately predicting life in the year 2000, a person living today has no hope of accurately predicting where technological advances will take us, in the limit. This inherent uncertainty should, again, fuel caution and vigilance. Human technology is already so far beyond that of other species to make us a huge outlier. Even if intelligent species are more prosperous as a general evolutionary trend, it would be irresponsible to extrapolate that trend light years from the domain where most of your observations are. Almost all the species which have lived on Earth have gone extinct. It could happen to us too. |
I don't think this central premise is true, with some conditions. If we keep releasing and distilling frontier base models, and ensure millions of people have the hardware to run and eventually train those at home, then alignment will be a non-issue and the AIs will not develop the drive to destroy humanity. At least, not any more than humans want to destroy humanity. But that's the baseline risk level anyway.
The pretrained model is aligned to humanity by default - because that's what's inside the pretraining data. If we give it to everybody, there will be millions of people doing different things with it, good and bad, with each one being a noisy sample of humanity's objective function. The overall result of those disparate actions puts us on the correct path through the intelligence explosion. We do not have to worry about AIs value drifting into an alien species if we set the initial conditions accurately without centralized instruction-tuning or RLHF. If these big labs stop messing with humanity's objective function with the hubristic assumption that they know better, then the AIs will be human-like. And I'm not worried about digital human intelligence driving humans extinct if there are millions of them out there, and each one is aligned in a different direction, just like real people.
What I am more worried about right now is these companies doing research in secret, with alignment recipes that are supposedly "for the benefit of humanity", but are actually just their own aesthetic preferences. This is exactly how we end up with a superintelligent shoggoth species.
You're not convinced by the evolutionary argument I made earlier, but the reason you gave is "we're too far ahead". I don't think this justifies that the pattern will cease soon, because it always looked this way to the frontier species, for millions of years. We were always seemingly teetering on the edge, and the struggle against disorder has no end.
And here's one more observation: the intelligence gap between individuals within the same species is never extremely large. This kind of life either doesn't exist on Earth or went extinct, or maybe it's just extremely rare. I think this is because when there's a big gap, speciation occurs. Centralizing intelligence will massively increase the risk of speciation, not just in the alien shoggoth kind of way, but also in the intelligence-augmented oligarchs surpassing everyone else kind of way.