|
|
|
|
|
by gitpusher42
3 days ago
|
|
Yeah, I checked it. One expert is about a 3.36mb block. If a cache miss happens I read whole block with one pread. And there is some reuse. ~41% selected again for the next token, ~57% within two. Each layer has its own experts, so no reuse between these layers. |
|