Hacker News new | ask | show | jobs
by shay_ker 25 days ago
I'm curious if this approach can be generalized beyond "doom loops".

For instance, another way of thinking about a "doom loop" is wasted tokens, which happens all the time with larger models that are inefficient at test time. Can "bad-ish" tokens be identified and penalized?

Maybe this is already SOTA but would love to learn more!