|
|
|
|
|
by FrancescoMassa
14 days ago
|
|
For sure that’s absolutely true because you’re switching context from a kV cache to another one. We introduced 2 algorythms for solving this :
1) sticky : switches model only when convenient
2) smartsqueeze : fast advanced compression for reingesting context |
|