|
I have some sympathy to geohot's view when it comes to pure informational chatbots. It's a first amendment issue, I'm allowed to write and read books that are useful to getting away with crimes, etc. This obviously doesn't work at all when the agents start doing real things in the real world, though. "Hey AI, I don't like my neighbor, find an exploit in the firmware for his car and make the cruise control malfunction and crash him next time he gets on the highway". This is committing a crime, not just talking about theoretical crimes. The AI can and should refuse it. I think he's anticipating and discarding this objection with his introduction, which otherwise feels disconnected from the rest of the article. FWIW, I have changed a bike tire and I'm pretty sure most of the MTS at the big labs could. This sort of "they're just bookworms who don't understand the physical world" rhetoric aside, we are currently seeing a ton of effort and expense go towards giving the AI agents hooks into being able to perform as many real-world-consequential actions as possible. And you can do a surprising amount with just bits, from writing code to breaking into systems to sending some combination of emails, phone calls, and currency to instruct meatspace humans to do things, etc. |
This leads to LLMs refusing to do security work for anyone, because it can’t tell if it’s being done for good or evil purposes.
Which is precisely what we’re going through right now with frontier models and it’s terrible.
These proposals always assume some perfect mechanism for identifying the thought crime with triggering on the normal requests. The people who want to commit crimes just persist until they jailbreak the guardrails while the rest of us suffer with denials.
Also if you can’t imagine these guardrails being used by governments to control inconvenient speech, you probably need to think a little harder about the realities of how these will be used.