Hacker News new | ask | show | jobs
Making models smarter and cheaper at the same time
2 points by Wetime 9 days ago
https://arxiv.org/abs/2607.14431
1 comments

Byte-exact KV grafting stores verified reasoning on disk. Replaying it lets a frozen 12B LLM beat 31B models at 8,700x less energy.