Hacker News new | ask | show | jobs
by causal 5 days ago
So this shows distance relative to other models, but I don't have a good sense for what these numbers say in absolute terms.

K3-to-Fable is blue at 0.42. Is 0.42 meaningful, or did we set 0.4 as the lower bound because it makes 0.42 look significant?

Sol-to-Fable is 0.69. It's dark yellow, making this look VERY different from 0.42. But is it? What do these numbers mean in absolute terms?

2 comments

I also suspect the questions asked matter a lot, and the system prompts matter a lot, because "the map is built from nothing but the words they choose" - so this is more a measure of linguistic style than anything.

If you use the Claude Code harness on two models you will probably get very similarly styled output. I would not be surprised if K3 "stole" a lot of the harness that was leaked.

Notice how the only rows that even come close to being gray is GPT-5.4+Mini, diverging even from other GPT models. Is this because it has a wholly different training set, or (more likely IMO), did it just have a system prompt that leads to a different style?

I think it needs some more explanation - Fable is 0.42 from itself apparently so Kimi K3 is basically indistinguishable?