Y
Hacker News
new
|
ask
|
show
|
jobs
by
radialstub
1 day ago
Yes but that's because of the scaling laws for transistors. ML models seem to get better the bigger they are. If you want to compress the world's information, you need to look at all the information in the world.
1 comments
trollbridge
1 day ago
Qwen 3.6 27b is way ahead of Llama 70b.
link