Hacker News new | ask | show | jobs
by refulgentis 16 days ago
There’s a good eval floating around somewhere and tl;dr they’re awesome but the benchmarks are cooked, you’re better off with Qwen 8B Q4 than 27B 1b or ternary.

Thanks for being skeptical, I maintain a llama.cpp-based client and it’s frustrating how high expectations are for local AI bc the median effort level means people mostly assemble their expectations and understanding via marketing soundbites