Hacker News new | ask | show | jobs
by gpjanik 6 days ago
TLDR: the experiment asks for in-distribution responses and gets those.

The right answer here is to ask a LLM to create a scene similar in quality to those, but completely out of distribution.

I asked GPT 5.6 Sol to give me a pelican playing football on San Siro while smoking a cigarette, in AC Milan's t-shirt. While this sounds like higher complexity of a problem, the generations from current models often include additional details like scene composition, scarf, etc., I don't ask for, so I wanted to see what here is memorization vs. composition skill.

"write svg code of a fish playing football on san siro in ac milan's t shirt, with raybans on and a cigarette."

Try that on GPT 5.6 Sol, Fable, or whatever other model. It's chaos.