Hacker News new | ask | show | jobs
by thebigship 3 hours ago
I think this one has advantages over the “pelican riding a bicycle” one because it hinges on an anatomical feature that many models associate with royalty, “habsburg” being a lineage and “habsburg jaw” being an anatomical feature.

Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.

Mistral returned byte-identical output across separate calls.

Gemini narrates its work in 65 comments; Llama says nothing.

If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.

2 comments

The identical pair from Mistral took me off guard. Many of the other models were so varied between the runs which is more what I would expect.
I wonder if the setup accidentally hit a cache at some layer.
Bite-identical?
clearly I missed an amazing copy opportunity, thank you haha