Hacker News new | ask | show | jobs
by ponyous 5 days ago
I agree with the conclusion and am happy to see this blog post, but this killed a bit of credibility for me:

> Using a single LLM judge for scoring. Every score here comes from one model, GPT-5.6 Luna, looking at one image at a time. I didn’t do much alignment and didn’t check how often it agrees with itself on a re-run.

Having used a similar setup (with previous gen LLMs) to evaluate the 3D models that my product[0] generates, it turned out there was no correlation at all. LLM judgments were very much random and I assume judging SVGs is not that far from judging 3D models. I guess I have to re-test this with current gen.

[0]: https://grandpacad.com