Hacker News new | ask | show | jobs
by stusmall 6 days ago
I'm glad someone ran the numbers on this. Every single Simon Willison post of an SVG is followed with someone dismissing it saying "I'm sure they train on it by now." This is despite a good blog post with sound logic on how easy that is to catch. [1] Glad to see someone took the time for a quantitative analysis of dumb little animals riding dumb little bikes.

1. https://simonwillison.net/2025/Nov/13/training-for-pelicans-...

4 comments

I don't think this small amount generalization to other animals and vehicles is strong evidence they haven't trained on this, either directly or more generally.
Honest question how could they possibly train on this as there are no good SVG pelicans to train off right? So they’re just training off a bunch of bad ones which should lead to just bad pelicans, but the pelicans are getting better.
Training on generating SVGs directly at all is already fairly niche. Generating full scenes with a cartoony character is even nicher. But there's plenty of non-pelican cartoony SVG content out there (created, not written, by humans with vector design tools), and more importantly, plenty of vision models to give feedback on the output (just raster as a png). You could easily hill climb this niche skill, if you cared.
It’s not hard for a visual model to score the quality of that output though, which would be a pretty good fitness function.
I did that. This works to some extent but why bother. Sft brings you 99,9% there already. Rl helps more with syntax errors. Svg is code
They can easily afford 1000 human made pelican svg files if they want. I think you underestimate how much resources SotA AI companies have.

(I'm not saying that they did that. I'm just saying they can.)

I agree with the conclusion and am happy to see this blog post, but this killed a bit of credibility for me:

> Using a single LLM judge for scoring. Every score here comes from one model, GPT-5.6 Luna, looking at one image at a time. I didn’t do much alignment and didn’t check how often it agrees with itself on a re-run.

Having used a similar setup (with previous gen LLMs) to evaluate the 3D models that my product[0] generates, it turned out there was no correlation at all. LLM judgments were very much random and I assume judging SVGs is not that far from judging 3D models. I guess I have to re-test this with current gen.

[0]: https://grandpacad.com

There's that version of the argument version, but there's also the softer version: that there used to be no training material of illustrated pelicans on bicycles, but now you have actual artistically talented individuals drawing it and that could improve the performance even though the AI labs are sucking it up no differently than everything else.

This post proves that hasn't happened yet, either. Although maybe the bad results posted online are being trained on and that explains the UNDER performance.

If I were running one of these mega AI companies I would set aside a tiny team to produce and release a pelican model, explicitly trained for this.

The very best svg pelican on a bile generation model. Just for laughs.