Hacker News new | ask | show | jobs
by charcircuit 18 days ago
This article doesn't address the inference side cost. Not all tokens cost the same. If the response is predictable you get a few output tokens for for the price of 1. The further an output token is the cost of generating it grows linearly due to attention.