It's so absurdly sensitive. It bailed out earlier today working on a TypeScript client for a sensor network API which happens to include some temperature and pH sensors for tanks, which yes, are used for biology experiments. But wow, we're degrees of separation from the actual biology work.
It's making it very hard to justify even trying to use Fable. When it works, awesome; it's legitimately good. But I can't trust it to do a task without deferring to Opus and that's really annoying at times. I want to know what I'm getting up front, not after the fact.
It refused to give me plant care instructions for an ornamental sold at my local Home Depot because it decided it was highly invasive and dangerous to grow in my region.
In their defence, google searching anything about plants these days leads to awful results. It’s saturated with slop. A response from an LLM might be slightly more reasoned and targeted. It’s hard to tell. This is a category of knowledge that’s being destroyed by people gaming google and dumping huge amounts of bad LLM and image generation onto the web.
This is also true for pretty much any other category of search that you do as a layman. Specific, targeted queries for official documentation or research are fine, but if you look up basic information about cars, health, or basic computer troubleshooting a massive portion of the results are AI-generated.
If I'm going to be getting AI-generated results, I'd rather read the output of a model that has all the context of my specific situation than slop generated in bulk with a cheap model from six months ago.
I asked whether an outdoors mosquito trap product (via a screenshot) would negatively impact other insect species in my garden and it refused. Though quick internet search did reveal that it would harm and trap many other species of harmless insects.
I asked him about sharks to be able to answer my kids question and it got triggered somehow. Then again when I asked it if my code had bugs or vulnerabilities before I commit.
At some point just kill the thing, it's not able to work properly as it is.
I'm writing a programming language with a "capability security model". That's enough to trigger Fable, it won't work on the language. It's hilarious. The mere presence of the word "security" seems to be enough to trip it up.
Anthropic refuses to allow Fable to code review my interpreter's memory safety. It was funny at first, then it became disappointing, then insulting, and finally utterly infuriating because I remembered the fact I'm actually paying for this nonsense.
Cancelled my subscription today. Hope OpenAI isn't patronizing like Anthropic. I don't want to hear about their "safety" bullshit ever again.
Yes it has completely turned me around - was all in on Anthropic but now it just looks too risky. Better off leaning into open models. Even if I found a way to work with the restrictions as they are, who is to say they won't suddenly change tomorrow. It's not worth it.
I mean it's a fucking joke, I kept getting refusals on a code base I wasn't familiar with and it was literally just because there are some vars named DNA. Just absolutely stupid.
It's making it very hard to justify even trying to use Fable. When it works, awesome; it's legitimately good. But I can't trust it to do a task without deferring to Opus and that's really annoying at times. I want to know what I'm getting up front, not after the fact.