Hacker News new | ask | show | jobs
by chermi 22 days ago
Late to this post, but my impression was that later models would be more efficient per task? Wouldn't they save compute released fable 5, maybe capping the effort, if it is actually a better model?
1 comments

I'm totally feeling you here as well. Fable is definitely quicker to first output token on `claude.ai` (ie less internal reasoning tokens being generated), and given how much more expensive decode is than prefill, I'm sure that must pay for itself pretty nicely, on top of any architectural changes that they must have made since the opus 4 architecture was locked in.