Hacker News new | ask | show | jobs
by xquce 33 days ago
Surely the same can be said for the people saying the opposite?
1 comments

I didn’t make a claim. The parent explicitly said it was a misconception that inference is not profitable.

No one knows if it’s profitable or not so we’re left to speculate.

You can make a fairly decent assumption by calculating the margin on serving glm 5.2, and adding say 30% extra costs and it still leaves a healthy margin
Where are you getting 30% from
It was a rough heuristic for how much more opus/5.5 would presumably cost if you extrapolate from glm5.2 prices. In any case input tokens for 5.4/4.6 are 70-100% more expensive, cached about 3-100%, and output tokens anywhere from 60%-240% more as per all their current api pricing. I highly doubt 5.4/4.6 are that much more expensive to serve given how cheap and commoditized inference has become, and how comparable they are perfomance wise.