Hacker News new | ask | show | jobs
by Bnjoroge 32 days ago
You can make a fairly decent assumption by calculating the margin on serving glm 5.2, and adding say 30% extra costs and it still leaves a healthy margin
1 comments

Where are you getting 30% from
It was a rough heuristic for how much more opus/5.5 would presumably cost if you extrapolate from glm5.2 prices. In any case input tokens for 5.4/4.6 are 70-100% more expensive, cached about 3-100%, and output tokens anywhere from 60%-240% more as per all their current api pricing. I highly doubt 5.4/4.6 are that much more expensive to serve given how cheap and commoditized inference has become, and how comparable they are perfomance wise.