Hacker News new | ask | show | jobs
by mncharity 10 days ago
> the hidden state for different layers carry meaningful self-awareness signal for various situations.

Is it plausible to wonder if some developer judgement feels, like maybe "the code I just wrote is clean/crufty", or "things came together smoothly/janky", might have extractable signals in some models?

If so, might one create a shopping list of desired signals to check for in a model, as with activation steering concepts, where one checks whether and how hard each concept can usefully be nudged?

1 comments

YOu are thinking along the right direction, we are going deeper into the signals.