Hacker News new | ask | show | jobs
by GaggiX 18 hours ago
I believe OP posted it because the new Kimi K3 has 69 KDA layers (the rest are 24 Gated MLA), I think previous large Kimi models had only MLA layers.
1 comments

It's not the same KDA as used in Kimi Linear, though.
What's the difference? They are both called Kimi Delta Attention.
The differences are explained in section 2.1.1 of the Kimi K3 technical report: https://arxiv.org/pdf/2607.24653#page=4