Hacker News new | ask | show | jobs
by TurdF3rguson 5 days ago
The cost of those output tokens is not zero or marginal.
2 comments

You could run the inference locally using the open weight models. Then you'd still get to use the models without sending money overseas
The first token on a new AI rig costs $X (full capex cost) then every token after that costs virtually zero. Over time the cost/token trends towards zero (modulo opex). That said, AI does have higher opex than general SaaS so it can’t get as close to zero.

But that’s kind of a different question, the running of some service. The “product”, the model, is a collection of files. The “manufacturing” required to add another customer is “send them the files” and has ~zero marginal cost.

There is a finite number of tokens that a rig will turn out over it's lifetime. Divide the cost of the rig plus electricity by that number and you have your cost per token. Yes providers can screw up on scheduling and end up paying more than they should, but that's not magic.
The AI rigs cost money, sure. But no one is talking about regulating those. We are talking about regulating models, which do have zero cost of reproduction.

It’s like banning the leaked DeCSS key. Good luck.

Well, no. We're mostly talking about banning access to models running on someone else's infra.

I don't think many people in this thread have the resources to run a frontier model in their living room.

How do you do that without banning the model itself.

It doesn't matter who can or can't run it.

They are trying to ban something that can be represented entirely by a sequence of bytes. An illegal number, like the DeCSS key.

The fact that money is involved does not change the free speech aspects. This isn't CSAM or torture porn. It's data.

It's not the model weights people are panicking over losing access to though. It's the compute.