Hacker News new | ask | show | jobs
by HarHarVeryFunny 29 days ago
AFAIK Alibaba's Qwen3-Next architecture doesn't do anything different specifically for the middle layers - it uses a 3:1 mix of recurrent and regular attention blocks throughout the full depth of the transformer.

The motivation for this mostly recurrent hybrid attention is to efficiently support long context lengths.