Hacker News new | ask | show | jobs
by parsimo2010 14 days ago
They aren’t training at all. They are quantizing existing models, it’s just that the process is different. The 27B uses Qwen3.6 27B as the base model.