|
|
|
|
|
by einsteinx2
21 days ago
|
|
> I also did an N=1 test with the same prompt doing a large non-trivial change to the codebase (migrating from Sqlite3 to Postgres) with both Fable Medium and Opus Ultracode, then had a new Fable session compare the two PRs...it decided Opus’s was much better! I can link a Gist with the review if anyone is interested, but I can't share the code as it's a private repo. I really figured Fable would bias to favor its own code, but I guess not. And Opus costed less (in tokens and subscription limits) and took roughly the same time (though you can’t really measure time since it depends entirely on how many GPUs Anthropic allocates at that moment which constantly fluctuates due to usage, plus Fable seemed to have been getting way more allocation than Opus during this test period as Opus was running unusually slow all weekend while Fable was ripping though tokens). Haha I just gave the exact same prompt to Opus Ultracode and it thought Fable’s was better. Obviously this isn’t the most scientific test due to LLM non determinism, and I still need to manually review both to make my own decision, but the fact at least they seem to basically be a wash is pretty telling about how much of an improvement Fable is when you actually compare them as close to apples to apples as possible (aka similar actual effort/token spend/sub agent activity) |
|