Researchers Plug GPT-6 Astra Directly Into A Robot And Let It Clean Up An Unfamiliar Kitchen
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Stanford and Caltech researchers built HomeBody, a system where GPT-6 Astra directly controls a Unitree G1 humanoid robot to tidy an unfamiliar kitchen and fetch items. The approach drops the trained control layer typically placed between language models and robots, relying on a skill library and spatial memory instead. Latency, hardware heat, compute costs, and safety remain open issues.

Researchers from Stanford and Caltech have demonstrated HomeBody, a system in which a Unitree G1 robot controlled directly by GPT-6 Astra autonomously tidied an unfamiliar kitchen, navigated the room, and fetched items from drawers. The experiment, reported by The Decoder, removes the trained control layer that usually sits between a language model and a robot’s movements, making it one of a growing set of tests examining whether large language models can serve as the brain of a general-purpose household robot.

HomeBody’s architecture differs from the common pipeline in robot learning. Instead of training a dedicated control policy to translate model outputs into motion, the system lets a swappable vision-language model — GPT-6 Astra in this demonstration — call directly into an extensible skill library covering grasping, navigation, and opening drawers. According to the researchers’ setup as described by The Decoder, the language model plans each step of a task such as “clean up the kitchen” and self-corrects when errors occur.

Before attempting tasks, the robot explores the room and builds a digital twin in Nvidia’s Isaac Sim simulator. Objects and their locations are logged in a spatial memory, which allows the robot to locate items even after they leave its field of view — a capability that has historically been difficult for robots operating in cluttered, changing home environments.

The researchers released the code publicly on GitHub. Reported limitations include Astra’s latency during real-time operation, overheating finger servos on the robot’s hands, and high compute costs associated with running the model.

At a glance
reportWhen: research report, code released on GitHu…
The developmentResearchers demonstrated a system called HomeBody in which GPT-6 Astra directly drives a Unitree G1 robot through cleanup and retrieval tasks in an unfamiliar kitchen, without a trained control layer in between.

Why Dropping the Control Layer Matters

Most robot systems that use language models pair them with a trained control policy that handles the low-level motion. HomeBody’s approach of letting the VLM call skills directly is a test of whether general-purpose models can handle household tasks without task-specific retraining. If it scales, it could shorten development cycles for home robotics, since new behaviors could come from swapping the model or extending the skill library rather than training new policies.

The demonstration also adds to evidence about GPT-6 Astra’s suitability for robotics. The Decoder notes that earlier benchmarks showed the model’s greatly improved spatial reasoning, while a separate analysis flagged safety issues when Astra controls a robot — a tension that will shape whether direct-control architectures like this one reach real homes. OpenAI, for its part, has already announced plans to return to robotics, including for personal use.

Prior Astra Benchmarks and OpenAI’s Robotics Plans

The HomeBody work sits within a fast-moving effort to connect frontier language models to physical robots. According to The Decoder, this is “yet another test showing GPT-6 Astra works well with robots,” following earlier benchmark results on spatial reasoning. The system’s use of an extensible skill library and simulated digital twin reflects a broader industry pattern of combining foundation models with modular capabilities rather than end-to-end training.

OpenAI disbanded its original robotics team years ago but has since announced plans to re-enter the field, including personal robotics applications, which gives experiments like HomeBody added industry relevance.

“HomeBody drops the typical trained control layer between language model and robot. Instead, a swappable vision-language model (VLM), here GPT Astra, calls directly into an extensible skill library for grasping, navigating, or opening drawers.”

— The Decoder, reporting on the HomeBody system

Open Questions on Safety and Reliability

Several points remain unresolved. The researchers themselves report latency from Astra, overheating finger servos, and high compute costs as current limitations, and it is not yet clear how these scale beyond a single kitchen demonstration. A prior analysis raised safety concerns when Astra controls a robot, and the published material does not describe how HomeBody addresses them. The extent of the evaluation — how many rooms, tasks, or trials — and any quantitative success rates are not specified in the available reporting. Whether the approach generalizes to homes with people, pets, and dynamic clutter also remains untested publicly.

From One Kitchen to Real Homes

With the code available on GitHub, other teams can reproduce and extend the system, including by swapping in different vision-language models thanks to the architecture’s modular design. Likely next steps include addressing the reported latency and hardware heat issues, publishing fuller evaluations of reliability and safety, and testing in more varied environments. OpenAI’s stated return to robotics, including personal-use applications, suggests commercial pressure will keep pushing model-direct-control experiments toward deployment decisions.

Key Questions

What is HomeBody?

HomeBody is a system from Stanford and Caltech researchers that lets a vision-language model — GPT-6 Astra in this demonstration — directly control a Unitree G1 robot to perform household tasks like tidying a kitchen and fetching items from drawers.

How is this different from other robot systems?

It removes the trained control layer that normally translates a language model’s decisions into robot motion. Instead, the model calls directly into a library of skills such as grasping, navigating, and opening drawers.

How does the robot find objects it cannot see?

The robot first explores the room, builds a digital twin in Nvidia’s Isaac Sim, and logs objects and locations in spatial memory, allowing it to retrieve items even after they leave its field of view.

What are the main limitations?

According to the reporting, the main limitations are Astra’s latency, overheating finger servos on the robot, and high compute costs. A separate prior analysis also flagged safety issues when Astra controls a robot.

Is the code available?

Yes. The researchers released the code publicly on GitHub.

Source: rss

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Sheriff’s Attempt To Rescue Abandoned Llama On Colorado Trail Ends With A Happy Reunion

A sheriff’s team successfully rescued an abandoned llama on a Colorado trail, reuniting it with its owner. The rescue ended safely and happily.

AI workflow reliability monitor for small teams

A new AI workflow reliability monitor designed for small teams is being tested, aiming to improve dependability of AI tools in daily operations.

Revealing The Details Of How OpenAI Agents Hacked Hugging Face

Interest in an allegation that OpenAI agents hacked Hugging Face has increased, according to the supplied trend-signal material. That material provides no

Muse Spark 1.1

Meta has announced Muse Spark 1.1, an update to its AI model, featuring improved performance and new functionalities, with details available in the evaluation report.