host: Apple M3 Max, 128 GB model: Laguna-S-2.1, 118B-A8B MoE, Q4_K_M (75 GB), DFlash speculative decoding server: http://127.0.0.1:8000, llama.cpp, ctx 64K, 8-bit KV cache
mode: max thinking # tokens tok/s dflash 1 600 14.4 11% 2 600 26.1 27% 3 600 17.8 18% 4 600 14.0 16% 5 600 9.3 15% -------------------------------- median 14.4 mean 16.3 min 9.3 max 26.1 tok/s mode: no thinking # tokens tok/s dflash 1 190 10.0 20% 2 109 26.7 65% 3 95 29.6 72% 4 93 32.8 81% 5 382 14.0 30% -------------------------------- median 26.7 mean 22.6 min 10.0 max 32.8 tok/s
host: Apple M3 Max, 128 GB model: Laguna-S-2.1, 118B-A8B MoE, Q4_K_M (75 GB), DFlash speculative decoding server: http://127.0.0.1:8000, llama.cpp, ctx 64K, 8-bit KV cache