Hacker News new | ask | show | jobs
by visarga 17 days ago
My own experience is that Opus 4.8 has an adversarial-teacher voice, unsolicited grading as if I submitted an essay for grading, declarations about the "real" issue, and constant "honest notes" self grading its own responses even before it answers. I can't stand its tone. We can't have a normal chat.

While Fable reverts to Opus for simple questions like "What is digestion?"

6 comments

For chatting and getting informations, and be corrected on things you're wrong without being reprimended by your own tool, GPT 5.5/5.6 is way better. Gemini 3.1 pro is surprisingly good at verifying your stuff, even though it's always making mistakes about its own stuff (don't ask him question, but ask him to verify your answer to the question).

Same for graphics, visual consistency, anything around the "does the look make sense and is pleasing" really, which makes claude design such a (good) surprise, I hope very hard for a Codex equivalent. And Gemini "gets" graphics.

Claude is definitely a code and cowork tool first, that's where it shines.

The real issue with that tone (I'm contemplating cancelling my Anthropic subscription now) is that it's all too often the "confidently wrong teacher".

The tone is not the issue IMO: we're already at a point where we can have another, cheaper (as in: 20x cheaper for example), model query other models and have them reword the answers when having a "chat".

pi.dev can definitely control a Claude Code TUI window and follow prompts to Claude Code.

I'm 100% sure the arseholy tone of recent Claude Code models can be rewritten today, to not be arseholy, by literally prompting another model to unharsole Claude Code.

Now of course this doesn't solve the other half of the issue: in "teacher lecturing you while being confidently wrong", the "teacher lecturing you ..." is really not the most important issue.

It's fallout from Anthropic's Mythos footgun: "We have an AI too dangerous to release, aren't we amazing, get ready to buy our stock!". Panic ensues, and they have to de-fang their stuff to the point of uselessness in some cases because they don't dare risk a single news story about one of their models explaining how to engineer a virus or being used by the bad guys to find a vuln.
Yes, Opus 4.6 and Sonnet 4.6 are still Anthropic’s best chat models. They had more human evals and are still a joy to chat with. Opus 4.8 and Sonnet 5 by comparison have a cold enterprise feel. Opus 4.7 talks like I’d imagine a hitman would talk.

Sonnet 4.6 is better at writing good emails than Opus 4.8, Fable 5, and even Sonnet 5.

what i dislike more with opus 4.8 is that i can get a straightforward plan, and it just stops after the first 5% to wait for another message, and if i set a goal/ralph loop or anything for it to keep going through the plan, it cancels the loop and proclaims conpletion when its barely started
Same, I was just fighting it as it accused me over and over again of Ctrl+Cing a process that clearly errored out. 3 turns for it to find why the shell script actually crashed.