Hacker News new | ask | show | jobs
by nolist_policy 448 days ago
On the other hand ollama supports iSWA for Gemma 3 while llama.cpp doesn't. iSWA reduces kv cache size to 1/6.
1 comments

What’s iSWA? Can’t find any reference online
Gemma 3 has some layers with a context size of 1024 tokens and others having full length. You need to read the Gemma technical report.
interleaved sliding window attention