Y
Hacker News
new
|
ask
|
show
|
jobs
by
upboundspiral
38 days ago
you can run Qwen 3.6 35B-A3B (3 billion active parameters + some GB for context) that can easily fit into 10 GB of ram while the not currently active experts are offloaded to the cpu ram with llama-cpp.