Hacker News new | ask | show | jobs
by Synthetic7346 13 days ago
What models do you use for automated pen tests? Don't the Frontier labs block security stuff?
1 comments

Opus was pretty decent at that before the whole Fable drama.

In fact, for our work, which is absolutely not rocket science most of the time, we see little upside in using Fable. The output is still "good enough," but the token burnout rate is through the roof. For many tasks, we stick to Opus, or even Sonnet, and the output is just fine.

Fun fact, we recently prepared an AI recruitment task. The basic idea was that there are conflicting goals in the context. We assumed that AI would lead candidates to a dead end, and they'd have to figure out what was happening. We abandoned the idea as even Sonnet was handling it fine enough. And it burned way fewer tokens than Fable would.