Hacker News new | ask | show | jobs
by anonymars 1 day ago
I disagree, as it seems that we are confronting problems that result precisely from emulating human behavior, in all its unsound, ill-defined, and abstract "glory"

For example I'm not convinced we can solve prompt injection by technical means (filtering) any more than we can phishing. And if you accept that premise, perhaps it turns out that it's best to mitigate it in similar ways, by assuming at least one person (or agent) will fall for it and ensuring you can limit the blast radius no matter what

As in the allegory of the junior developer who deletes the production database: the fault lies with the fact that the developer could delete it

1 comments

Yeah, I don’t disagree with that, it’s a good framing and analogy. I thought you meant more the philosophical aspects. However some LLM behavior are also really not human like, for example no human would panic failing to make the business profitable after 23h, to the point where they need to start doing crazy stuff. I would expect a human to just give up