Hacker News new | ask | show | jobs
by jmugan 2 hours ago
A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
4 comments

Aren't they still bad at understanding how bicycle frame works? Especially the steering part?
Shhh...you'll alert the models :)

Totally agree though, anyone with a vague understanding of how bikes works ignores the pelican because they know the bike is unrideable in the first place.

Yes, but I think the idea here is that most models produce very similar pelicans on bicycles, so a different test might be more useful in gauging the differences in models.
Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting.

A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.

But it's another benchmark on how good models are at generating intensely average, unwanted things with unthinking design. Just scaled up.
Bad? It has a charming style. I would watch the whole book if it was made like this.
yeah, definitely, in the same way that we all regularly go and look back fondly at our chatgpt ghiblified family photos