Hacker News new | ask | show | jobs
by kamranjon 7 days ago
Whoa whoa whoa, 118b params, 8b active MOE, long context reasoning, open weights - music to my ears. Hadn't heard of this lab before but I am very excited, will definitely try this out tomorrow - this is a real sweet spot I think in terms of model size and performance.
1 comments

If the numbers are legitimate then our prayers have been heard