|
|
|
|
|
by HarHarVeryFunny
29 days ago
|
|
AFAIK Alibaba's Qwen3-Next architecture doesn't do anything different specifically for the middle layers - it uses a 3:1 mix of recurrent and regular attention blocks throughout the full depth of the transformer. The motivation for this mostly recurrent hybrid attention is to efficiently support long context lengths. |
|