Hacker News new | ask | show | jobs
by senordevnyc 17 days ago
I don’t know, I can imagine that the equivalent of a PIP for an agent would be noticing flaws in its output and then creating additional evals, modifying the harness, upgrading the model, etc. to improve its performance. That’s not SO different from the PIP, except that the LLM doesn’t really care in the same way?
1 comments

The person making the tweaks to the harness is the one taking accountability for fixing those mistakes - they're the DRI in this scenario.
Like the manager being responsible for someone on a PIP.

Obviously they’re different, but it’s interesting to think about what is unique about humans that makes them able to be accountable or responsible in a way that LLMs cannot be. Is it at core just that they can be fired? Or essentially the threat of suffering?

I think just the possibility of consequences that they can give a damn about. An LLM doesn't have feelings no matter how much human-like text it can simulate. There is literally nothing going on between prompts for any given model. They are incapable of worry or any other emotion. Wipe the context clean and the LLM is completely unaware there was ever a problem.