|
|
|
|
|
by toasty228
2 days ago
|
|
> absolutely no way to stop it. Unless if they take down the flat rate subs and you have to pay api prices, usage will drop a lot. Right now I freelance using gpt, if I had to pay api prices I'd probably lose 50% of the income, if not more, with the ~40% tax on top I might as well spend my time doing something else |
|
To give you an example of model in this class, the DeepSeek V4 Flash preview is a 284B A13B model, with a native quantization mixing 4 bit and 8-bit values. You can easily run it on an RTX Pro 6000 Blackwell (or 2) at a reasonable quant, especially if you offload the MoE weights to 48-64GB of system RAM. This costs US$11,800 at Microcenter right now, and it will work in any gaming box with decent cooling and a modern 1000W power supply. Over the lifetime of the card, an entire system would cost you under $4,000/year. Power is about 300W for the card (either a blower model, or a workstation model with the power cap), and another 150W or so for the rest of the server.
Or you could buy it on Open Router from dozens of different commodity vendors, starting around $0.09 per million tokens input, $0.18 per million tokens output. This is a competitive market price, so some of the providers might be losing money or reselling surplus capacity. But given the underlying hardware costs, the numbers are in the ballpark. In other words, if you're willing to settle for lower-quality tokens, you can get roughly Sonnet 4.5 for close to free, or as a modest capital expense for a successful freelancer. Halfway decent tokens are cheap, and you can generate them in-house!