|
|
|
|
|
by x312
14 days ago
|
|
It's been known for several years that LLM activations encode future tokens ahead of time (e.g. https://arxiv.org/abs/2404.00859). But this has only been shown on simple tasks, so I think this paper is still quite neat. The interesting thing is that they show "future horizon length" varies across models. |
|
Of course, an interesting question what part of this internal computation is modeling for the future compared to guessing based on the given context (the past).