Karpathy’s Pelican
https://twitter.com/karpathy/status/2083749667410727319Spaghetti 2026:
https://x.com/dreamingtulpa/status/2083304533829066873
https://xcancel.com/dreamingtulpa/status/2083304533829066873
That method of perception probably scales N^2... so sure with more compute, LoTR animation will improve. But I think to get a real jump in "experiential feedback", perception needs to scale linear or sublinear. Maybe that's there LeCunn's jepa will come in.
There needs to be the removal of the middle man:
image -> text -> action
To image -> action.
Here's minecraft: https://threejs.org/examples/webgl_geometry_minecraft.html
Here's an FPS: https://threejs.org/examples/games_fps.html
The library is extremely well-documented. When Three.js vibe coded projects started blowing up on Twitter 1-2 years ago, I wasn't that impressed then either because I knew what it was doing.
Anyone who remembers the C compiler built by an LLM! backlash probably feels the same way:
Why would I use an LLM to create a well-known demo rather than fork that demo itself?
What I haven't seen yet from an LLM is it create anything fundamentally new and exciting: A new UIX that is actually good. A new service that is actually good. A game with an art style I haven't seen, music, storytelling - something you'd expect out of a AAA studio.
Given that rant: The coolest part is the multi-modality between text and animation. However, I think the end product would have been a lot better if it was just a video. Having it do it in Three.js didn't add a ton of wow factor for me, and it would have been a lot better looking as video.
> Something like an ephemeral GTA of X on demand.
Here's where you lose me. AAA gaming is very far away from this Three.js demo. But the novel part being the syncing of narrative to the visual scene - a text-to-audio (video) book type technology seems very possible (and useful).
Nice work. We need more big projects like this involving AI (if anything, to get away from the slop argument largely focusing on 1-shot experiments).