|
|
|
|
|
by orbital-decay
5 days ago
|
|
>character trigrams Worthless. Any LLM output similarity metric that uses n-grams as a source might as well measure the average temperature on Mars surface, no matter how much lipstick you put on it. It just can't have enough certainty. There used to be an n-gram benchmark popular on Twitter that showed extreme similarity of grok-3-beta to gpt-4.5-preview, while these models were trained on new base ones, came out 2 weeks apart, and were unmistakably different. Results were wildly inconsistent run to run. It didn't stop the crowd believing its creator in that DeepSeek R1 was trained on o1-preview (which was obvious bullshit as well, they were as different as two models can be). It's amazing how you can put anything on the web and everybody will believe you without checking or even understanding of what they're looking at. K3 was trained on Claude's outputs, though - it repeats Anthropic's prompt injections 1:1 in its reasoning, which you should know if you ever tinkered with both models long enough. Good for them. |
|