Hacker News new | ask | show | jobs
by mdasen 12 days ago
It also depends on how many tokens it needs to burn through to accomplish something.

At this point, I always look at things like Artificial Analysis' total cost to run their tests. It'll take into consideration the cost of tokens, how many tokens it burns through, and how effectively it uses caching (and the price of that caching).

If a model "costs the same" but its reasoning ends up going through a ton more tokens, it doesn't really cost the same in real world usage.

1 comments

Precisely. GLM 5.2 Thinking is pretty damn good - but it regularly does something nonsensical. Or even spits out what looks like a fragment of its memory cache. Or returns a bunch of Chinese.

I find myself having to resubmit a query very often...so it being a third of the cost of other AIs isn't really relevant.