Spaghetti 2026:
https://x.com/dreamingtulpa/status/2083304533829066873
https://xcancel.com/dreamingtulpa/status/2083304533829066873
A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.
I have always preferred the result of getting an LLM to draw an svg or make a procedural animation like this to the uncanny hyper-realistic result of diffusion image/video generation.
But I do think this was a poor demonstration of the idea for another reason: Tolkien works have a HUGE corpus of training data. It's great that random users can come in and immediately recognize what the footage is, but it fails at the very first thing the pelican was meant to do:
- Render this thing you have only tangential training data of, that also happens to be an asymmetrical object so we can see how much you fuck up the details if you somehow flip the orientation half the time.
They should have used an obscure story, not "Most Studied Piece of Literally Work of The Past Century trademarksymbol"