|
|
|
|
|
by sroerick
2 days ago
|
|
Not to mention - If you switch the view to "coding tasks" on this website: Kimi K3: $3.18 per task
GLM 5.2: $6.51 per task
GPT 5.6 Sol: $7.02 per task
Opus 5: 8.23 per task
Fable: 11.70 per task
So it's pretty dang cheap lol. Nobody is using frontier inference for "office tasks". |
|
General office work is one of the big frontiers the labs are pushing on, and it's part of how they're justifying the value proposition to enterprise customers. It's also accounts for a big portion of the spend on RL; tasks/environments designed to train agents to navigate Slack or Salesforce. If you're Anthropic pitching Claude to a bank (taking an example I'm familiar with), coding probably accounts for ~20% tops of the workforce, and it doesn't drive direct revenues. The 'agentic coding bump', but for all your analysts, traders, and wealth managers, would be a much more attractive prospect.
I don't disagree that coding is the most successful use case so far (and probably more relevant to a HN audience). But I think the future of the labs is also contingent on them making progress on more general white collar work. I suspect that's why the Opus 5 release blog lists 3 coding benchmarks (FrontierBench, DeepSWE and FrontierCode) to 3 or 4 more general ones applicable to office work - depending on how you slice it (GDPVal, AutomationBench, Legal Agent Benchmark, BrowseComp).