|
|
|
|
|
by sidkshatriya
461 days ago
|
|
As per the technical report, every 5 layers you have a global attention layer. The global attention layer during training can have as many as a 128k context length during training (though I understand it is usually 32k). Q. When you are training with a context length of 128k, is the attention in the global layers dense or sparse ? If dense, would the attention memory requirement here would be O(n^2) where n is 128k for each global layer ? |
|
We wanted the long context recipe to be friendly for finetuning, and training at 128k is a bit of a pain we don't do it. For inference, we see inference at 128k with the 5/1 is close to RAM usage for a fully-global-layer model at 32k.
Individual attention layers are always dense.