Hacker News new | ask | show | jobs
by evansjp 13 days ago
hey that's awesome! yeah the eval showed first pass was only ~65% real decisions. The fix that stuck was an entry has to name a real file it touches or it gets dropped. A code decision names code.

I agree agents don't always self-talk decisions, that's why we distill the whole transcript after the fact instead of asking them to log anything. Your baselining idea is good!