|
|
|
|
|
by mallowdram
267 days ago
|
|
There are no world models in there, it's trained on arbitrary images/sequences.
There are no world models in us, we learn from only specifics in topological space, stitched together in sharp wave ripples. Everything is from detached memories working through optic flow. That's not a world model, it's not even a model. It's an analog.
This whole world model thing is another branding phase after language models failed to deliver. After world models it will be neuro symbolic, then RL will sweep in like a final boss fight, and then... it still won't work. Notice anything about these names? They're walking pneumonia paradoxes. |
|
But what's coming is: Vision-language-action models and planning, spatial AI (SLAM with semantics and 3D reconstruction with interactability and affordance detection). Video diffusion models, photo-to-gaussian-splats, video-to-3D (e.g. from Hunyuan), the whole DUSt3R/VGGT line of works, V-JEPA 2 etc. Or if you want product names, Gemini Robotics 1.5, Genie 3, etc. The field is progressing incredibly fast. Humanoid robots are progressing fast. Robotic hands with haptic sensors are more dexterous than ever. It's starting to work. We are only seeing the first glimpses of course.