Hacker News new | ask | show | jobs
by solenoid0937 26 days ago
Fable always felt clearly a huge step above Opus for me. It's been able to one shot complex bugs and apps Opus could never solve. But it's expensive.
3 comments

Honest question/comment for you and the parent: I find these subjective experience reports pretty empty without an understanding of your level of experience, the problem space you're working in, etc.
I think the improvement on how it codes is pretty much represented correctly by the benchmarks (a nice bump, but not some crazy leap)

But where it really shines is in how NOT lazy it is. Fable requires less hand-holding. And I can understand how someone who uses Claude-Code sparingly and with very focused prompts would not see a lot of improvement there.

But simple example: if you ask Opus to do a review of the codebase (with a short prompt and not too much guidance), I've had it basically read the `git log` output, do a simple `ls` and have it declare "Everything looks great! No problems found!", when Fable really does what you would expect it to do.

And you might think: "oh, so it's just capable of handling crap prompts?", well sure. But even if you make THE PERFECT Opus plan (a plan that would take many turns/hours to finish), Opus will fake out, say everything is done, and then you see that half of the plan was deferred, half of the functions are ridiculous stubs, ...

If you give the same plan to Fable, it'll just DO IT. And it WILL get it done. And in the end it'll tell you "Oh, I also found 30 other bugs and I fixed all of them properly" (where Opus would have started crying, or WORSE, worked around the bugs)

> Opus will fake out, say everything is done, and then you see that half of the plan was deferred, half of the functions are ridiculous stubs, ...

Doesn't Claude Code have a /loop command? Give it a message to keep it on track overnight, send every 20m, make it track progress in a doc, reread the doc after every loop. I've found this works well for a certain class of problems, most importantly where the actual work is getting done by very narrowly focused batches of subagents, with the main session just coordinating and keeping the doc updated.

They added a "/goal" command which I guess spawns a supervisor agent process that checks to see if your goal statement has been achieved (e.g. "/goal complete tasks 1-250 of plan.md") I've been pretty happy with it but I rarely use that workflow. Most of the time I give it a 3-6 step prompt and come back in 20 min and the first two were done and I get a summary "up next is to complete the next steps" which.... Opus 4.6 didn't have this problem. 4.8 feels like a cost cutting measure, or maybe it's just tuned poorly for my specific workflow (multi-repo system integration)
I tried using `/goal` when it just came out, but back then they used Haiku for it. And if your main model is a 1M model, Haiku can't even read that much. So my /goal always failed. (I instead went for an elaborate /loop-scheduled message)
For optimizations or proofs I suppose? Wouldn't know why else you would do something like that.
I think the parent comment stands - I’ve asked Opus to do a review of DeepSeek’s test suite and told it a couple things I wanted it to look for, and it did a very thorough review of the tests and picked out a reasonable number of gaps and tautological tests. It’s a mix of prompting/instructions, the agent harness, and random chance. The model is not wholly irrelevant but IMO increasingly so.
20 yoe, application/systems stuff, and I always run models on xhigh or max effort level.

Fable has been more intelligent, with better taste and defaults (e.g. make impossible states impossible without being told, build for testability), and considers/solves things that Opus did not.

My workflow is to run Claude in planning mode first to spit out a plan file and then review->revise cycle it with Codex or other agents.

One big tell is that Opus will say that it can't find any more revision advice for a plan file, yet Fable will find more issues but also smart pivots into better solutions. This is probably the best test since it's not based on vibes.

15+YOE. Fable 5 is well above the level of Opus. I have used it alongside Opus for a range of hard problems, including porting a large static analysis tool to Rust, building various tooling around .pptx and .xlsx documents.

In all cases, Fable clearly outperformed Opus.

I'm doing work with fairly complicated cryptographic algorithms and math. I'm finding Fable 5 to be a significant stop better than Opus 4.8, but that Opus occasionally comes up with something small but nontrivial that Fable missed. (The reverse is true much more often.)
That's the delta in our use cases then, I suppose. I'm not doing anything super novel. DevOps work, web application development — things that typically do not stump the agent(s) when given time to iterate.
Yeah, I've learned that it's only worth deploying Fable for the most challenging problems. For a while, my Fable workflow was looking like ths:

Me: Hey Fable, I've got this massive, theoretically challenging, totally novel, ill-defined cutting-edge problem that I'd like you to solve.

Fable: < Doesn't merely solve the problem -- utterly obliterates it. Nukes it from orbit. Does a robust one-shot that takes several hours to complete. >

Me: Holy smokes, that was amazing!!!! But the formatting could use some simple refinements. Could you change the margins and maybe add a drop-cap at the start of each section in the user docs?

Fable: < Commences another multi-hour nuclear exchange with the code >

Me: WT?!?!

(The moral of this story is that bringing a nuke to a knife-fight is only occasionally the best strategy. And in more practical terms: Fable is amazing -- but only for certain classes of problems, and even if it were free there's a lot I probably wouldn't use it for.

~13 yoe, and I had some nasty WebRTC + CallKit problems that Opus couldn't make a dent on but Fable figured out.
What is your view on how experience and problem space relate to subjective experience.

For example will inexperienced or experienced users see a bigger jump in subjective quality?

The main difference I'd guess is whether your prompts are targeted or broad.

Less experienced people tend to use very broad prompts.

Experienced people tend to understand the structure of the code and give explicit guidance such that a larger model isn't necessary to read between the lines.

I noticed with GPT-5.6 (through work), I could step up my specificity by a level of abstraction. But I still intentionally scope the prompts fairly tightly, as I find it produces better results if you need to own and maintain the code.

It still does stupid stuff like leave unnecessary abstractions around after refactoring instead of proactively suggesting to remove them.
Amusingly, I was impressed with Fable's puissance at coding in one particular session, shortly after they turned it back on. True to its reputation, it displayed an accomplished mastery of the problem domain and relentlessness at refining and testing the solution I asked for.

Then I checked /usage and discovered I was still running Opus 4.8 xhigh.

Only version week-one.

I’m downgrading tomorrow.

It’s horrible slow and it feels like opus very often. It’s a totally different experience from the first week