Hacker News new | ask | show | jobs
by ainch 2 days ago
Agreed, Kimi is cheaper for coding - I say that explicitly in the post too. However I'd have to disagree with you on the "office task" front.

General office work is one of the big frontiers the labs are pushing on, and it's part of how they're justifying the value proposition to enterprise customers. It's also accounts for a big portion of the spend on RL; tasks/environments designed to train agents to navigate Slack or Salesforce. If you're Anthropic pitching Claude to a bank (taking an example I'm familiar with), coding probably accounts for ~20% tops of the workforce, and it doesn't drive direct revenues. The 'agentic coding bump', but for all your analysts, traders, and wealth managers, would be a much more attractive prospect.

I don't disagree that coding is the most successful use case so far (and probably more relevant to a HN audience). But I think the future of the labs is also contingent on them making progress on more general white collar work. I suspect that's why the Opus 5 release blog lists 3 coding benchmarks (FrontierBench, DeepSWE and FrontierCode) to 3 or 4 more general ones applicable to office work - depending on how you slice it (GDPVal, AutomationBench, Legal Agent Benchmark, BrowseComp).

1 comments

Okay, I think that's fair, but I'm not convinced there's anybody actually doing large amounts of compute on office tasks? Do you know anybody? Can you point to anybody publicly documenting this? Can you can you point to any specific workflows where fable is being used in lieu of more basic models?

Even if you provide exceptional answers for all of these I still think it is disingenuous at best to ignore coding tasks in writing this. I have to assume coding is 90% of the use cases for the frontier.

You've made a case that the labs need these customers. You haven't made a case that the labs have these customers.

No I think you're right that the amount of compute spent on office work is lower than coding - although I don't have any sense for the right share. The best source I could find was an OpenAI report [1] which mentions that ~64% of enterprise token generation is via Codex, which I would expect to skew entirely towards coding. But it's hard to say how the remainder is split, what proportion is 'frontier', or whether it's representative for Anthropic.

On your questions - I've spoken to a number of execs and seniors behind closed doors but nothing public I can point to. Anecdotally, I've spoken to senior leaders at banks spending billions of tokens on one-off tasks like prepping execs for earnings calls or piloting end-to-end agent workflows for specific use cases (but mostly piecemeal/one-off).

On Fable, financial analysts I know are using it to produce research docs, models and decks - I hear that it's a big improvement for these tasks. This lot have been blindsided by the spend growth [2], the same as for coders in enterprise (e.g. Uber blowing annual budget in 4 months [3]), so I do think there's appetite and budget for a capable, cheaper open model - but, due to the price, Kimi does not obviously fill that role the way it might for coding. That said, I still largely agree with you on share - where coding has seen a broad deployment across software development, most of the office work stuff is still fairly piecemeal and certainly lower compute-spend.

I think it's fair to say I could've focussed on coding more rather than taking AA's benchmark distribution as representative - perhaps a more balanced title would be "Kimi K3 is not cheap across the board"? I guess there's also some ambiguity about what 'cheap' means - as I said elsewhere in this thread, I think when some people talk about the price of Chinese models, they imagine Deepseek competing with o1 for 1/20th of the price. Even though it is better priced for coding, Kimi isn't Deepseek-level cheap.

I do, however, think you could debate whether coding will remain at >50% total token usage going forwards - big enterprises are hunting for ways to get value out of LLMs, and the labs are investing a correspondingly large amount in generating demonstrations and RL environments to get the models up to par. At the end of the day, programmers make up ~5% of all white collar work. Of course, it's also possible that Chinese labs will shift focus to white collar applications now they've demonstrated a lead on coding cost efficiency, so, I mean who knows - it'll be interesting to get some detail when Anthropic IPOs.

Sorry for the long reply! Appreciate it's quite meandering...

[1] https://cdn.openai.com/pdf/5d1e1489-21c0-43e4-9d42-f87efdbf0...

[2] https://www.reuters.com/business/finance/australias-cba-flag...

[3] https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-c...