Y
Hacker News
new
|
ask
|
show
|
jobs
by
nolist_policy
448 days ago
On the other hand ollama supports iSWA for Gemma 3 while llama.cpp doesn't. iSWA reduces kv cache size to 1/6.
1 comments
vlovich123
447 days ago
What’s iSWA? Can’t find any reference online
link
imtringued
447 days ago
Gemma 3 has some layers with a context size of 1024 tokens and others having full length. You need to read the Gemma technical report.
link
nolist_policy
447 days ago
interleaved sliding window attention
link