Hacker News new | ask | show | jobs
by _mrinalwadhwa_ 10 days ago
Was able to run it on Apple M3 Max (128 GB)

host: Apple M3 Max, 128 GB model: Laguna-S-2.1, 118B-A8B MoE, Q4_K_M (75 GB), DFlash speculative decoding server: http://127.0.0.1:8000, llama.cpp, ctx 64K, 8-bit KV cache

  mode: max thinking
   #  tokens   tok/s  dflash
   1     600    14.4     11%
   2     600    26.1     27%
   3     600    17.8     18%
   4     600    14.0     16%
   5     600     9.3     15%
  --------------------------------
  median  14.4   mean  16.3   min   9.3   max  26.1   tok/s

  mode: no thinking
   #  tokens   tok/s  dflash
   1     190    10.0     20%
   2     109    26.7     65%
   3      95    29.6     72%
   4      93    32.8     81%
   5     382    14.0     30%
  --------------------------------
  median  26.7   mean  22.6   min  10.0   max  32.8   tok/s