Hacker News new | ask | show | jobs
by ianbutler 13 days ago
> But long-context is a trap, because performance still falls dramatically after 150k-200k context.

I often see this repeated, and it is not true task to task. I work on this daily and we have several tasks where long context is advantageous and our evals against a whole battery of models with different windows show it as being so.

This is why having good evals for the tasks you're working on is so important.

I do grant it's a good rule of thumb.

1 comments

That might be me getting paranoid but I actually see it above 200k (with Opus 4.8), but more often and MUCH more pronounced in the afternoon UK time.