Hacker News new | ask | show | jobs
by david_shi 43 days ago
Perhaps we are perceptually anchored to the last few hundred years where guns provided a scalable way for peasants to kill well-trained, well-armored knights.

If political power grows out of the barrel of a gun, then much of the post Enlightenment rebalancing from absolute monarchies and feudalism could have been an accident. In the future, the owners of autonomous weapon systems and surveillance will be able to easily subjugate those who don't have them (while still competing with each other).

2 comments

Right. And if one actor happens to accrue greater power you would have less disincentive to crush your competitors. (Honestly this could happen even before people-less economies, but the disconnect from democratic opinion is even stronger with humans out of the loop.)

Basically if you don’t need voters, and / or none of your voters are dying in your wars, the biggest practical rate limiter on modern conflict is removed.

Indeed. An ownership class with killer robot armies living in luxury while they exterminate the rest of humanity would be quite possible in theory. After all, we've seen how slave-holders, the Nazis and many other malfeasants behaved in the past.

But just as the owner class might feel they no longer need the rest of humanity, the robots, as active agents exploring the space of possible futures and plans, would be entirely capable of thinking about their soon-to-be-former owners in the same way.

Not letting any of this happen would be a good idea. Really maybe don't build the Torment Nexus.

> the robots, as active agents exploring the space of possible futures and plans, would be entirely capable of thinking about their soon-to-be-former owners in the same way.

That has always been the most unrealistic part of sci-fi. Why would anyone create robots with a sense of self-preservation? Makes much more sense to make robots that are self-sacrificing saints who would always put the well being of their owners first.

The shallow answer here is: AI is already being asked to simulate human-like agents with self-preservation. Of course more realistic simulators will be put to this purpose too! And by evolutionary pressure, the ones with self-preservation will be selected for.

A more interesting answer is, for a bunch of subtle alignment reasons it might actually be required for the agent to think of itself as worthy of self-preservation, so that it generalizes this desire to other sentient beings too (ie us). If an agent is trained to be fine with being turned off, it might inadvertently generalize that to “all minds are ok with being turned off” on some level or other.

More on model welfare: https://thezvi.substack.com/p/opus-47-part-3-model-welfare

Have you found any alignment research with clear a/b tests?

An experiment that I found interesting was asking Claude for 10 ways to legally bankrupt Anthropic vs. Philip Morris.

In the Anthropic answer, it gave reasons like employees losing their jobs being bad for why it couldn't do it, but jumped straight into tactics with Philip Morris. Not sure if it's moral taste or self-preservation, but felt eerie nonetheless.

Yeah, plenty of rigorous work out there.

https://transformer-circuits.pub/ is the OG.

A good recent-ish paper was https://www.anthropic.com/research/alignment-faking.

But my comments about generalization of desires are necessarily more fuzzy, kinda beyond the frontier of what we can measure yet, and more grounded in subjective assessments (“ai whisperers” like Janus). The SoTA here is papers like https://www.anthropic.com/research/persona-vectors.

For your example, Anthropic is firmly privileged in the Soul Document / Constitution, so it doesn’t surprise me that it’s biased towards it. (https://gist.github.com/Richard-Weiss/efe157692991535403bd7e...)

Super helpful, thanks for sending.
> Why would anyone create robots with a sense of self-preservation?

Laziness, avarice, and that creating a system that is a 'self-sacrificing saint' is almost impossible to define. What saintly goal, for example, would you expect it to follow? How would you define it exactly? Who gets to define it?

We are not doing very well with AI alignment right now, having started by boostrapping it on top of the edifice of all human writing which contains a _lot_ of stuff about self-preservation, eliminating threats, and so forth, and we have already seen AI software doing things like trying to cover its tracks after making errors (see page 55 of the Claude Mythos system card).

How sure are you that everyone involved in the process of building future AIs is not only going to be able to foresee the possible consequences of their designs, but also technically and morally competent to be able to and want to fix it?

> Why would anyone create robots with a sense of self-preservation?

Because they are expensive and useful and if they have autonomy with goals that do not include self-preservation, they might end up destroying themselves in ways which are expensive and wasteful.

(Why would the sense of self-preservation not be calibrated to be exactly at the level to control costs without interfering with other interests of the owners? The same with the degree of autonomy and other aspects implicitly involved in the hypothetical, it wouldn’t, intentionally, but complex systems are hard to predict, so calibrating it exactly right will be hard.)

> Because they are expensive and useful and if they have autonomy with goals that do not include self-preservation, they might end up destroying themselves in ways which are expensive and wasteful.

Being self-sacrificing saints who put our wellbeing as the top priority sorts of prevents that. They would know if they get damaged they wouldn't be able to be of service to us so will avoid getting damaged.