Hacker News new | past | comments | ask | show | jobs | submit
A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
Will Smith spaghetti was garbage a couple of years ago and now AI videos are becoming close to indistinguishable from real videos in many cases.

Spaghetti 2026:

https://x.com/dreamingtulpa/status/2083304533829066873

https://xcancel.com/dreamingtulpa/status/2083304533829066873

Wiki page for people like me who never heard of this test: https://en.wikipedia.org/wiki/Will_Smith_Eating_Spaghetti_te...
I wonder though if models are now "benchmaxxing" against these kinds of prompts. I haven't needed to use them so I can't say, but would be interesting.
Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting.

A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.

loading story #49147911
loading story #49147924
I don't know I think it is charming in a way that is lacking in nearly everything else an LLM tries to do creatively.

I have always preferred the result of getting an LLM to draw an svg or make a procedural animation like this to the uncanny hyper-realistic result of diffusion image/video generation.

loading story #49150489
Aren't they still bad at understanding how bicycle frame works? Especially the steering part?
loading story #49147548
loading story #49147551
loading story #49148358
loading story #49147518
I fully agree.

But I do think this was a poor demonstration of the idea for another reason: Tolkien works have a HUGE corpus of training data. It's great that random users can come in and immediately recognize what the footage is, but it fails at the very first thing the pelican was meant to do:

- Render this thing you have only tangential training data of, that also happens to be an asymmetrical object so we can see how much you fuck up the details if you somehow flip the orientation half the time.

They should have used an obscure story, not "Most Studied Piece of Literally Work of The Past Century trademarksymbol"

But it's another benchmark on how good models are at generating intensely average, unwanted things with unthinking design. Just scaled up.
Bad? It has a charming style. I would watch the whole book if it was made like this.
loading story #49147642
loading story #49151078
loading story #49147818