Hacker News new | ask | show | jobs
by korrectional 20 days ago
My opinion is that claude code uses more tokens simply because Anthropic makes more money that way and forces people into their subscriptions. This is supported by the fact that they won't let you use your sub on a different coding agent. I use pi btw.
9 comments

Once I realized that Anthropic is a token merchant, I start to understand Anthropic’s decision more. They are always finding reasons for you to use more tokens through them unless the users revolt or demand some guardrails.
I've done a couple side by sides on web chat with the same prompt on Opus 4.6, 4.7, and 4.8 and the output gets longer/more verbose on version increment. The enerr variants are definitely much wordier.

On the other hand, the newer variants also tend to benchmark higher so it's not quite a clean argument of "hey the new version eats more tokens"

I think both things can be true: new models benchmark higher and eat more tokens.
From my experience new models are slower and use more tokens even on questions which gpt 4 answered correctly. It is mostly because newer models tend to be more verbose (even with prompt requesting short answers).
Unless somebody improved on the underlying transformer architecture... Surely AI is smart enough to do it by now
I've done a couple side by sides on web chat with the same prompt on local 4b, 14b, 32b open models and the output gets longer/more verbose on version increment.

Its rather frustrating, slower tokens and more tokens.

I bailed on Anthropic the moment they started blocking alternative harnesses like pi on their subscription plans.
If I were anthropic I’d force that too. They offer the harness and if they control the entire pipeline then they can optimize the entire experience. It doesn’t have to be nefarious.
This is kind of a strange comment as it implies a false dichotomy.

Its not 'nefarious' in that its in their best business interests.

But it'd be difficult to take anyone serious who thinks Anthropic's motivation was to improve the UX, and the other effect were by accident. At the time they specifically started blocking based on openclaw prompt text. Its a walled-garden tactic.

A walled garden is nefarious to people who do not want to be inside one.

It's like Microsoft banning Vim users that use Azure
They didn't ban people from using Claude, though. They banned them from their flat-fee subscription and required that you pay per token.

It's still questionable but I don't think it's in the same ballpark as what you describe.

I don't think it's in the same ballpark at all. I checked the `/usage` in my session which uses a Max x5 plan. One day I had used $400 of tokens and 20% of my Fable allocation. Anthropic is effectively giving us more tokens per $ on the monthly plans but it comes at the cost of Anthropic being the prompt-writers and managers of the agents pretty much entirely. I don't think this is a bad deal.
It’s really not. Vim isn’t instrumental to Azure usage.
CC isnt instrumental to use Anthropic LLMs. Yet here we are.
> It doesn’t have to be nefarious.

The nefarious part is because it's non optional. They could give you an option and compete by being better, instead you're given the finger as the option is taken from you. Competition is hard and banning people to create more FUD serves business need better.

You've obviously been gaslit so badly you're desperate to find a way to defend a shitty move and pretend it's the only way to increase usability. But you don't have to deny really! You're allowed to admit control is easier for a company than competition, and that they didn't have to, but did because it increases their control of the ecosystem.

If you want to defend someone, good? But at least save it for someone who actually deserves it. They don't; and you insult you and your readers intelligence by trying.

> if they control the entire pipeline then they can optimize the entire experience

The only issue is that Anthropic optimizes the entire experience for their bottom line. User experience and price only suffer becaue of that.

> if they control the entire pipeline then they can optimize the entire experience

So what? When you care about optimising the entire experience, you offer sane defaults.

When you prevent people from changing the defaults, it's about control, not experience.

Sounds like they're modeling their PR on the classic Apple playbook: "choice is bad, and you should appreciate the constraints we've generously imposed"
The Agents are more like Double Agents. Purporting to work for you, but with the primary goal of siphoning your wallet to its handler.
But they gave us double the tokens! Then a limited time more usage! Then even more tokens "off peak" times! Then some new model released but apparently it inherently used 1.69x tokens! Then Fable is here but "it uses much more usage". But only until ~~the US banned it~~ ~~7th July~~ ~~19th July~~ who even knows.

At this point I think Dario is just in his wellness retreat adjusting a revenue/profit dial.

Ah, the ol' retail switcharoo.

Increase the price by 70% and then cut it by 50%, resulting in a 15% cut that sounds like a major deal.

now reealize that LLMs are trained to produce tokens and like the halting problem, cant be trained not to produce tokens and youll realiE the AI labs are the perfect essential capitalist and like cancer, will keep growing useless tokens until it kills its host.

no amount of alignment will stop aomeone drom just shutting up.

LLMs might be trained to produce tokens, but Anthropic don’t have to price by tokens. If an organization is a ‘non-profit’ and they decided to design their pricing to be tokens-based, I get it. If a for-profit design their pricing to be tokens-based, I don’t know where are they drawing the line between profit vs benefit. That doubts makes it hard for me to be a customer. Disclaimer, I still use Claude…
tokens definitely measure compute.
You can ask it to verbatim produce training data and that takes very little compute for a lot of output tokens
i dont think you understand how these models operate.
Serious Willy Wonka energy?
Seems unlikely they'd be this dumb. The way to get us to use more tokens is to make those tokens more useful, not less. Anthropic is full of people (including higher-ups) who know this.
But it is much much simpler to make it consume more tokens.

It’s like that saying “What Andy giveth, Bill taketh away”, but in this case it is one company.

There is definitely a conflict of interest.

It's the same conflict of interest quite literally any business has. What stops any business from over-charging? Competition.
> What stops any business from over-charging? Competition.

I fully agree.

> It's the same conflict of interest quite literally any business has.

I know that you know what I meant ;) In the long term it is just as you say - overcharging (eventually corrected by competition forces), but in the short term it can be additional revenue, blamed on a bug, but making some manager look good.

I thought I read somewhere that according to filings for going public, subscription revenue is tiny… like 5%.

Edit: consumer Claude subs are the 5%. I’d bet most all of CC subs lump in under enterprise.

  - API & Enterprise: 75% to 85% of total revenue.
  - Business Subscriptions: Roughly 10% to 15%.
  - Individual Subscriptions: About 5%.
So the incentive to have Claude Code use more tokens should be even stronger then as AI & Enterprise are using consumption based pricing.
The vast majority of my company's enterprise plan use is through Claude Code even though we have access to the API and could be using OpenCode instead.

I don't fully agree with the premise that they intentionally increase system prompts, but the enterprise plan usage is going to make that a huge income for Anthropic.

The fact that individuals are more likely to use the alternatives than businesses is telling.

Anthropic is fine, as long as someone else (a clueless employer drinking Dario Koolaid) is paying for it. But the moment you have to pay for it, people just bail and go for DeepSeek, Kimmi, OpenRouter, OpenCode Go and other alternatives that give more bang for the buck than Anthropic.

Yep. That's my case.

I now have unlimited Anthropic and OpenAI plans at work because CEOs bought the hype.

But for personal projects $10/mo OpenCode Go serves me with DeepSeek V4 Flash, MiMo 2.5 and GLM 5.2.

You're making the opposite argument. Anthropic is incentivized to use less tokens in Claude Code because people are paying a fixed monthly fee for subscriptions.
Nope, that’s not true, because they want you to pay for the higher subscription bracket.
Can confirm — they got me paying $100/mo this way.

Also I think it’s well known that OpenAI is the much less expensive option (in tokens and $$). For the same $20 you get a lot more mileage.

Curious if folks have strong opinions about the overall UX of OpenCode vs CC…

For me as well, at least this month to use more of Fable. We'll see if they extend Fable access because of people like me.
That strategy only makes sense if there's an abundance of tokens, but that's not the case. AI companies are spending a ton of resources on improving token efficiency because they are all severely GPU constrained. Anthropic instead nudges you to move to a higher tier by setting rate limits.
Also not true, they want you to pay for a higher subscription bracket and then use only marginally more than you would have, which I think they’re doing quite effectively for most people based on my interactions.
If they wanted to play games with sub tiers they would just change the rate limits rather than wasting inference.
Flip side is customer psychology. Choosing a more expensive tier leaves better emotion.

Also i doubt there was jira ticket with “make llm more verbose”, rather ticket with “bug makes llms too verbose” gets prioritised taking revenue impact into account.

Higher subscription brackets are likely worse for them. I recall seeing someone calculate that a fully maxed out highest subscription bracket is something like $15K in tokens?

And people paying $100 or $200 are much more likely to max it out for purely psychological reasons - it crosses that threshold where I want to see my money's worth in full. Whereas people on $20 subs are more likely to be there just to get access to better models and features, and are not necessarily even doing any substantial work.

It's always more complicated than that because the prestige and Early Adopter users are what drag other people to also be customers to avoid FOMO.

Your gym members who got a subscription aspirationally and don't show up are absolutely subsidizing the power lifter who is introducing wear on (tens of?) thousands of dollars of equipment three times a week, but if the regulars weren't there you wouldn't have sold those subscriptions at all. Without a poster child there's no poster.

They could just use less tokens and finish your quota sooner. So even tho I think are a bad company, I can’t say they do this for the reason you said.
Well since what you get for your subscription is unknown it would be trivial to get that result without burning tokens.

Especially since compute is such a scarce resource.

Generally, companies with >150 people can’t use subs. So yeah, it’s mostly a funnel for devs/small companies to eventually vet for the product and convince their enterprise to use it as well.
Enterprise users are not paying a fixed fee, though
Yeah, I strongly recommend against Claude Enterprise, it is ridiculously expensive and hard to control costs.
Not really. The incentive is to make you hooked on the process, so you bring the same process to the workplace, and start paying corporate prices, not individual subscription prices. For that to work Claude Code, prompt, and the rest of the mechanics has to be more or less uniform.
> I use pi btw.

When using Pi, one way to significantly reduce input tokens it yields is to ignore common bookkeeping "dot directories", such as `.git`. How to do so can be found with the following interactive Pi prompt:

  How do I configure Pi to ignore git related artifacts, such 
  as the project's .git directory?
Other local assets to consider ignoring are `.pi`, `.agents`, `*.md`, and language specific output directories such as `__pycache__`, `bin`, `obj`, `target`, etc.
> I use pi btw

Not sure if intentionally meant as a reference, but it gives "I use Arch btw" vibes.

Pi is one of the ways out of this problem (OpenCode another) so I took it as an intentional reference as it is highly relevant. I also use Pi as my daily driver and I think it's a wise choice to figure out how to decouple yourself from lab-specific harnesses that you have little control or observability over.
the amount of system prompt wastage going on in orgs is insane. we identified 400k in annual burn for zero value in just one section of our large company.

and the interesting thing about system prompt wastage is its a cost that scales non linearly with subagent use.

The non-linearity is interesting. Is the default behavior for subagents in CC/OpenCode loading the same full system prompt (or AGENTS.md)?
I'm sorry, what! 400k...?
> My opinion is that claude code uses more tokens simply because Anthropic makes more money that way and forces people into their subscriptions.

OTOH, this makes typical subscriptions usages consume more tokens, which are included in their flat fee.

This sounds more like incompetence than malice.

It would be true if there was a unified "Anthropic" entity making every decision from pure rationality. Instead, more tokens increase Claude Code team's metrics of token usage, which most likely has a KPI around token usage and adoption.

To remind Goodhart's law: "When a measure becomes a target, it ceases to be a good measure".

..also to parent's point, yes the upsell is only appealing once user run's out of tokens.

Anecdotally at least Claude code uses less api money for me than other harnesses. I think people might be missing some caching discount?
> This is supported by the fact that they won't let you use your sub on a different coding agent

I mean, that's a very weak argument? Isn't a much more plausible explanation that with your tooling you'll have more of a lock-in than with just your model?

Neither is mutually exclusive.

They get lock-in, and through that lock-in are more effectively able to inflate token usage.