|
|
|
|
|
by andai
19 hours ago
|
|
>You Could Have Come Up With Kimi Delta Attention What? Little old me! Well, then, let's have a look... > (First paragraph) > A note on notation: this article defaults to bra-ket notation because (in my quantum-inspired opinion) it makes the shapes in this derivation very clear. The Math notation switch above rewrites every equation using conventional bold vectors and explicit transposes instead. In bra-ket mode,
∣
q
⟩
∣q⟩ is a column vector,
⟨
k
∣
⟨k∣ is a row vector,
⟨
k
∣
q
⟩
⟨k∣q⟩ is a number, and
∣
v
⟩
⟨
k
∣
∣v⟩⟨k∣ is a matrix. Vectors face right by default, while keys face left when written into the linear-attention state. We work with one causal attention head and real-valued vectors, assume DeltaNet’s keys are normalized, and let the state map from key space to value space. Hmm... Guess not! |
|