|
|
|
|
|
by cyb3ralbert
13 days ago
|
|
Context size cuts like this are usually a cost/latency tradeoff rather than a capability one - serving a smaller window is cheaper and keeps latency in check, and most sessions probably don't need anywhere near 372k tokens anyway. Curious if this affects people who were actually relying on the larger window for big codebases. |
|
The issue was more specific to higher token burn rates, not latency.