Hacker News new | ask | show | jobs
by nextaccountic 25 days ago
Funnily enough, Roko's Basilisk might as well be a self-fulfilling prophecy: perhaps future AI models may be trained on texts about it and pick up traits consistent with torturing people that didn't help develop advanced AI

If nobody ever talked about it, I doubt any AI agent would think of this dumb idea on their own

... which may be a reason to ban talking about it

3 comments

That's actually an important part of the theory of Roko's Basilisk. The danger of being tortured only applies to those who are aware of it. Supposedly, the incentive to torture you only exists if you were aware of the implied threat of torture.
It's even dumber than that. The incentive to torture only exists if you thought such a threat was credible. Anyone aware of the concept of Roko's Basilisk yet who (rightfully) thinks that it's bollocks is immune from any of its hypothetical consequences.
Steelmanning, I think the argument is that a vindictive but fair AI would not torture someone who did not cooperate with it because they didn't believe the threat was real, because they were merely mistaken, not malicious. It's similar to how a just god would not damn sincere unbelievers, it would only damn true believers who nevertheless refuse to worship it.
Or the AI companies could filter it from their training data. That would be another, probably easier, option.
Schrödinger's basilisk?
You have to observe it to force the decision: slither away, or attack.