Hacker News new | ask | show | jobs
by bredren 11 days ago
I just tried this on the monthly $18 plan, having it do a basic task with its 2.7 model and then audit it using k3.

K3 got into some loop trying to run docker and after maybe the 6th attempt ran out of quota for the 5 hour window which represents 20% of the weekly.

I run 200 max and chatgpt pro, but I had to blink at that.

K3 didn't even write out what it was doing or provide any sense for why it was pursuing the execution path it was.

I'm in disbelief that this is a groundbreaking model, and do not think it represents a threat to Claude Code or Codex at this time.

1 comments

Have you used the same session for audit? so switched to K3? or used a new session for K3? K3 is sensitive to this, they wrote about it on their blog.
Do you have a link handy?

I used same session, set it to k3 model. I’ll look at the blog but the result was so bad I am prepared to abandon.

I should have saved the output.

I think maybe it was a mistake to not use open router.

K3 is very different to K2, I wouldn't be surprised if there are different system prompts, parsing templates, etc; which confuse/poision the model's context.
I can believe it, but if it is such a threat, why not warn or prevent it from harness level?

I'm remembering now warnings to this effect early on in CC.