How long did it take vLLM to implement deepseeks sparse attention from the r1 paper?
Does ananyone outside deepseek have a working code for the v4 compressed attention mechanism?
Has any other provider managed to bypass CUDA and program the compute engines in their native assembly language to get 10% more performance out of them?