Hacker News new | ask | show | jobs
by sailfast 22 days ago
How long will that $4.40 rate persist? Until we know more about the real unit economics it will be damn near impossible to rely on steady inference costs or make them predictable at the enterprise level. Gonna be a wild ride for awhile.
1 comments

Multiple providers (who need to make a profit) offer the same 4.40 rate for glm-5.2. It's not subsidized.

Deepseek's 0.86 or whatever is likely subsidized but alternate providers offer it for a price comparable to glm-5.2.

According to deepseek themselves, their current rates are NOT subsidised.

They have published tons of articles dedicated to performance and efficiency engineering. Feel free to have a look...

Why is no other inference provider offering similar prices then?
How long did it take vLLM to implement deepseeks sparse attention from the r1 paper?

Does ananyone outside deepseek have a working code for the v4 compressed attention mechanism?

Has any other provider managed to bypass CUDA and program the compute engines in their native assembly language to get 10% more performance out of them?

There is your answer.

GPU/RAM/etc prices could continue to rise. If the world leaders decide it's time to build the robot armies, then that could price out the civilian uses for GPUs.