|
|
|
|
|
by ponyous
5 days ago
|
|
I agree with the conclusion and am happy to see this blog post, but this killed a bit of credibility for me: > Using a single LLM judge for scoring. Every score here comes from one model, GPT-5.6 Luna, looking at one image at a time. I didn’t do much alignment and didn’t check how often it agrees with itself on a re-run. Having used a similar setup (with previous gen LLMs) to evaluate the 3D models that my product[0] generates, it turned out there was no correlation at all. LLM judgments were very much random and I assume judging SVGs is not that far from judging 3D models. I guess I have to re-test this with current gen. [0]: https://grandpacad.com |
|