|
|
|
|
|
by erwan577
18 days ago
|
|
The KV-cache memory usage also seems remarkably frugal, even at the full context length. That could make this model particularly useful in multi-agent coding workflows. I wish KV-cache memory usage and related optimizations were discussed more clearly in new model announcements and demos. |
|