Hacker News new | past | comments | ask | show | jobs | submit
I think the most important insight is the limitation of LLM perception:Slowly taking screenshots.

That method of perception probably scales N^2... so sure with more compute, LoTR animation will improve. But I think to get a real jump in "experiential feedback", perception needs to scale linear or sublinear. Maybe that's there LeCunn's jepa will come in.

There needs to be the removal of the middle man:

image -> text -> action

To image -> action.