Stanford and Caltech connect GPT Astra to a humanoid robot

Stanford and Caltech connect GPT Astra to a humanoid robot
News

Stanford University and Caltech have presented HomeBody, a humanoid-robot system that uses the frontier vision-language model GPT Astra to work through an unfamiliar kitchen. The September 2026 project page shows a Unitree G1 exploring the room, tidying several objects and retrieving an item from a drawer after receiving an underspecified request. The Decoder independently reported the demonstration on September 27, adding details about the system’s current limitations.

HomeBody is interesting because it questions a common robotics stack. Many systems put a large vision-language model in a planning role, a trained vision-language-action model underneath it and a low-level controller below that. The researchers instead give the language model persistent spatial memory and a library of reusable humanoid skills. GPT Astra can call abilities such as navigation, grasping and drawer opening, while the lower-level skills and controller still handle the physical execution. This is direct high-level orchestration, not a claim that a chatbot emits raw motor commands safely.

The robot first gathers information with onboard cameras and other sensors. The team combines those observations with SLAM, joint data and waypoints to build a digital twin in NVIDIA Isaac Sim. The resulting spatial memory can record where objects are and what has happened. In the kitchen demonstration, the robot makes multiple trips to pick up coffee bags, discard spoiled milk or orange juice and place items where they belong. A second task asks it to find medicine in a drawer and throw away the carton. The system uses feedback from its actions to continue or correct the plan.

The setup is still a research prototype. The official page notes latency when using Astra remotely and describes a local stack that currently needs a powerful RTX 4090 laptop. The Decoder also reports overheating finger servos and high compute costs. Stanford’s public repository says the full code is still coming soon, and the demonstration does not establish reliability across homes, objects or safety-critical tasks.

The broader lesson is about physical AI rather than a ready-made household robot. A capable general model may coordinate a changing environment through a stable skill library, reducing the need for a separate policy trained for every room. That could make experimentation faster, but it shifts pressure onto memory accuracy, latency, hardware robustness and permission boundaries. HomeBody is therefore a concrete step toward more general robot agents, with the difficult engineering and safety work still ahead.