Hacker News new | ask | show | jobs
by therobots927 36 days ago
I didn’t make a claim. The parent explicitly said it was a misconception that inference is not profitable.

No one knows if it’s profitable or not so we’re left to speculate.

1 comments

You can make a fairly decent assumption by calculating the margin on serving glm 5.2, and adding say 30% extra costs and it still leaves a healthy margin
Where are you getting 30% from
It was a rough heuristic for how much more opus/5.5 would presumably cost if you extrapolate from glm5.2 prices. In any case input tokens for 5.4/4.6 are 70-100% more expensive, cached about 3-100%, and output tokens anywhere from 60%-240% more as per all their current api pricing. I highly doubt 5.4/4.6 are that much more expensive to serve given how cheap and commoditized inference has become, and how comparable they are perfomance wise.