Yea, looks like it (at least for their v2). Fable did much better in V1, but once they added their tooling around it, Opus + Composer ended up doing better (>0.8 Grade on the SQLite chart) for the final product. Granted, Fable reached a 'good' result (0.7~) in a much faster timeframe.
Point is, frontier models don't necessarily are better than a swarm of non-frontier ones for these scoped problems. IMO frotnier models excel at underspec'd or more complex problems where ideation and exploration are key (and that's where they become crazy expensive).
Anthropic wants it that way so they can’t be sued… the correct answer is don’t, switch to Kimi K3 for much more usage, the same quality of model, and no hidden practices :/