| > What worries me about this is that Anthropic and OpenAI seem to have backed themselves into a corner of high costs. Can they reasonably decrease their prices by 20-50x to compete with DeepSeek or Xiaomi’s Mimo? They have high prices, not high costs. They will obviously keep prices as high as they can for as long as they can, while keeping demand up. Once demand starts to fall, so will the prices. > Are these models cheap because they are open weight and having hundreds or people stress test running them on different hardware helped to lower the cost? Or is it that they are being provided as loss leaders to drive the prices down? Neither. They are cheap because they have neither technical edge nor brand power to keep the prices high, and so have to ask commodity prices for them. People somehow still don't get it, despite everyone who studies the economics of it telling them: Inference is dirt cheap. Training is expensive, inference is cheap, and getting cheaper. |
Money is made on the subset of inference that is charged at cost + margin via their APIs. API usage is so high because customers are still finding their feet, trying to understand how to measure the value they get from their spend, erring on the side of spend.
Yes, in a world of unmeasured value and tokenmaxxing, inference is profitable on SOTA models because all capacity is being consumed at all times, driving down marginal costs, but what about a world in which capacity isn’t constrained? There are still huge fixed costs.
Even the most optimistic leaks with the current high prices put the margin on API token inference at around 50%. How can SOTA models ever come close to competing on price? Price always matters. Offering the best model with the most brand recognition does not exempt OpenAI from the basic rules of business.
Historically, software has been such a successful business because the margins are incredible, 95%+ in many cases, driven by direct measurable value to customers that dwarfs the cost. A 50% margin at a time when your customers are falling over themselves to spend as much money as they can is not a good sign, it is a very bad sign, it leaves no room to ever achieve traditional technology margins, and inevitably leads to very weak margins.
Inference needs to become an order of magnitude cheaper than the value it delivers to ever have a chance of delivering on this wildly profitable vision. The cheap model providers have a much better chance of achieving that.
Outside of coding, almost every business case for AI doesn’t need above human intelligence, it doesn’t even need human intelligence, or half a human intelligence, a business can extract a lot of value from a machine that has a fraction of a human’s intelligence. Most human work does not use our intelligence, it is rote, a monkey could do it, and that’s where AI will be used most. Who is going to pay $10 per million tokens when they could pay $0.10 to get the same outcomes?