Helix 2.5 AI Model Introduced as Humanoid Robots Manage Household Tasks in 30 Unseen Residences


We independently review everything we recommend. When you buy through our links, we may earn a commission which is paid directly to our Australia-based writers, editors, and support staff. Thank you for your support!

  • Figure reveals the Helix 2.5 AI model with notable advancements in household generalization.
  • Helix 2.5 was evaluated in 30 unfamiliar residences, achieving a 56% success rate.
  • Physical Intelligence’s pi 0.5 model employs multimodal data for generalizing household tasks.
  • Robotic foundation models demonstrate promising advancements, yet practical implementation is still years away.

How Helix 2.5 addresses unfamiliar homes

Figure describes Helix 2.5 as its most advanced neural network to date. Rather than obtaining demonstrations within the specific homes where the robot would operate, Figure pre-trained the base model on Index, its global dataset of human behavior videos.

From that foundational model, Figure adapted three full-body actions: tidying a living room, folding towels, and making beds. The team then deployed its humanoid across 30 residential locations throughout the San Francisco Bay Area.

  • No training data was collected from any of the 30 test residences.
  • None of the evaluation items, towels, or bedding were present in the training data.
  • The robot utilized the existing beds, couches, and tables that each home had.
  • Grading was strictly binary, requiring complete end-to-end execution with no partial credit.

The contrast in performance between policies with and without foundational pre-training was pronounced. When Figure tested a control policy trained from scratch without Index pre-training, it achieved a zero-shot success rate of merely 9 percent across the homes. However, when employing the Index pre-trained Helix 2.5 model, that success rate rose to 56 percent.

Figure also showcased an empirical human-to-robot transfer scaling law. By training four models with an eightfold increase in pre-training data while keeping downstream tuning constant, the action prediction error consistently decreased with every doubling of data. Figure asserts that the trend was consistent enough to predict the validation loss of its largest model run to four decimal points before training.

The Airbnb tactic and the divide to active family life

Viewing the footage released by Figure founder Brett Adcock reveals insightful details about the execution of these tests. Figure conducted evaluations across 30 properties, which appear to be holiday rentals and Airbnbs rather than occupied family homes.

Utilizing short-term rentals is an ingenious engineering shortcut. It provides the team with immediate access to diverse floor plans, varying mattress heights, and multiple surface textures without upending employee households.

This approach also underscores the disparity between an empty rental and true domestic life. The homes displayed in the demonstration videos are tidy, well-lit, and unoccupied. In several clips, company engineers can be seen closely observing the robot as it maneuvers around furniture.

Navigating static furniture in an empty rental without prior mapping is a significant accomplishment in autonomous spatial reasoning. However, an active Australian family household presents a considerably more chaotic setting.

In reality, dogs may dart through the kitchen, children can leave school bags scattered in entryways, and family members walk about. While static spatial generalization is a crucial first step, safely coexisting with people in dynamic environments remains a challenge yet to be addressed.

Physical Intelligence adopts a multimodal approach with pi 0.5

While Figure prioritizes human video pre-training for humanoid platforms, Physical Intelligence approaches generalization from a complementary perspective with pi 0.5.

Physical Intelligence constructs vision-language-action foundation models. The central concept behind pi 0.5 is heterogeneous co-training. Instead of training solely on actuator telemetry from a single robot platform, the model incorporates a mix of multimodal web data such as image captioning and visual question answering, along with action data from static dual-arm setups and mobile manipulators.

This dual-path strategy fosters semantic understanding alongside low-level actuator control. When given an instruction like “clean the bedroom,” pi 0.5 generates a high-level subtask in text, effectively communicating with itself to break the task into manageable steps. It subsequently funnels that step into a continuous flow of 300 million parameters, directing physical joints in one-second action segments.

Physical Intelligence evaluated pi 0.5 across unseen homes on domestic tasks, including placing dirty dishes in sinks, loading clothes into hampers, and wiping countertop surfaces with sponges.

Their ablation tests indicated that web-scale multimodal data was the primary contributor to the robot’s ability to recognize unfamiliar household items. Incorporating data from other robotic designs offered physical baseline stability. After training in around 100 diverse environments, pi 0.5 achieved generalization performance in new homes that closely matched baseline models trained directly within the target test rooms.

Maintaining perspective on progress

These announcements affirm that robotic foundation models are advancing swiftly, but it’s vital to keep timelines realistic. We are likely still years away from entering stores like JB Hi-Fi or Harvey Norman to purchase a domestic humanoid for A$15,000 to handle our Saturday cleaning, yet this suggests it isn’t a decade off.

A 56% zero-shot success rate in unfamiliar homes indicates a significant increase from 9 percent, yet it also signifies the robot fails more than four times out of ten. If a household appliance showed that failure rate, it would remain unused in storage. Humanoid hardware continues to be expensive, power-intensive, and mechanically intricate.

The significance of these technological updates lies in validating the software roadmap. For years, the industry contemplated whether physical manipulation could benefit from the foundational model scaling laws that propelled large language models.

The data from Figure and Physical Intelligence substantiate that physical intelligence scales with data. Training foundation models on a wide array of human activities and multimodal data equips machines with a fundamental intuitive comprehension of the physical world prior to entering a room.

We are still in the early stages of this journey, but the era of hand-coding every single room is finally nearing its end.

Summary

Figure’s Helix 2.5 and Physical Intelligence’s pi 0.5 models signify substantial advancements in AI robots managing household chores. While promising strides have been made, obstacles persist regarding practical implementation and attaining higher success rates in real-world scenarios.

Reader questions

Frequently asked questions

Fast answers to the questions readers ask most about Helix 2.5 AI Model Introduced as Humanoid Robots Manage Household Tasks in 30 Unseen Residences.

What is Helix 2.5?

Helix 2.5 is the newest AI model from Figure aimed at enabling humanoid robots to carry out household tasks without prior mapping.

How was Helix 2.5 assessed?

It was tested in 30 unfamiliar residences, achieving a 56% success rate in tasks such as tidying, folding, and making beds.

What is the pi 0.5 model?

The pi 0.5 model from Physical Intelligence employs multimodal data to understand and execute household tasks, presenting a different methodology for task generalization.

What challenges do AI robots face in homes?

Challenges consist of achieving higher success rates, adapting to dynamic environments, and lowering costs and complexity.

Posted by David Leane

David Leane is a Sydney-based Editor and audio engineer.

Leave a Reply

Your email address will not be published. Required fields are marked *