Brain · AI-derived
Figure AI Helix 2.5 Delivers 56% Zero-Shot Success Across 30 Homes
Figure AI released Helix 2.5 on September 17, 2026. The Index-pretrained model enabled Figure 03 humanoids to tidy rooms, fold towels, and make beds in 30 unseen Bay Area homes with no new data or fine-tuning, raising full-task success from 9% to 56%.
ZeroGantry analysis
At 56% full-task success after halving adaptation data and scaling to 30 homes, Helix 2.5 implies roughly 2x lower behavior development cost per deployed unit once fleets exceed a few dozen robots. Operators should watch continued Index scaling for reliability gains above 80% before committing to unsupervised home pilots; ignore only if per-home data collection remains cheaper than the $3.5B pretraining path for their use case.
Helix 2.5 and the Zero-Shot Generalization Milestone
On September 17, 2026, Figure AI announced Helix 2.5, a vision-language-action foundation model pretrained on its Index human-behavior dataset. The release detailed results from 420 trials in 30 rented Bay Area homes never encountered during training or adaptation. Three whole-body tasks—tidying living rooms by placing 13-15 toys into a basket, folding four towels into a basket, and making beds by positioning pillows and smoothing comforters—were executed using a single fixed checkpoint per behavior. No data collection, weight updates, or environment-specific fine-tuning occurred in the test homes.
The controlled comparison isolated Index pretraining as the variable. A policy trained from scratch achieved 9% full-task success under identical task data, architecture, and evaluation criteria. The Index-pretrained Helix 2.5 reached 56%, completing 237 of 420 trials with no partial credit. Per-task rates showed variation: bed making at 67% (94/140), towel folding at 62% (87/140), and tidying at 40% (56/140). This gap directly attributes the generalization lift to pretraining on diverse human video rather than task-specific data alone.
Figure also compared Helix 2.5 against a prior Helix 02 policy trained with data collected directly in its evaluation environment. Helix 2.5 matched that policy’s success rate while requiring only half the task-specific adaptation data and generalizing across 30 homes instead of one. Behavior specification costs dropped by a factor of two while the deployment scope expanded thirtyfold. These efficiency gains address a central economic barrier for household humanoid deployment, where per-home data collection has historically scaled poorly.
Index Dataset Pretraining and Scaling Laws
Index, introduced publicly on August 25, 2026, aggregates human behavior video captured via a mobile app network. At launch it reported 264,000 downloads across 108 countries, over 44,000 weekly active users, 16 million videos uploaded, and $15 million paid to contributors. Processing reached 30 minutes of video per second, equivalent to roughly 4.9 years of human activity daily. Per 1,000 hours collected, the dataset contained 373 unique tasks, 1,146 manipulated objects, and 116 environments after a five-stage pipeline of filtering, fraud review, deduplication, rebalancing, and annotation.
Helix 2.5 pretraining on Index produced a measurable human-to-robot transfer scaling law. Doubling pretraining data volume improved downstream robot-action prediction loss predictably enough for the team to forecast the largest training run’s final loss to four decimal places before execution. At the time of the Helix 2.5 announcement, Index generated approximately 35 minutes of new human experience data every second. Figure committed $3.5 billion in compute through Nscale, targeting up to 100,000 NVIDIA Vera Rubin GPUs for continued scaling.
No single evaluation task exceeded 1.90% of the Index pretraining distribution, ensuring broad coverage rather than narrow overfitting. This diversity supports the observed transfer: policies adapted from the same base model handled locomotion, rigid and deformable manipulation, bimanual coordination, and active perception across novel layouts and objects.
Implications for Household Humanoid Deployment
Current commercial and research humanoids typically require environment-specific data collection and fine-tuning before reliable operation in new settings. Helix 2.5’s results indicate that large-scale human-behavior pretraining can reduce or eliminate that step for a defined set of chores. Service providers could deploy fleets with a common base model, then adapt behaviors once using modest task data collected elsewhere, cutting per-site onboarding costs and time.
The 56% full-task success rate remains well below levels needed for unsupervised home service. Strict criteria—no partial credit, safety interventions counted as failures, and timeouts—highlight remaining failure modes in perception, grasp stability, and long-horizon planning under novel lighting, clutter, or object configurations. Bed making outperformed toy tidying, suggesting deformable object handling and precise spatial reasoning still vary by task complexity.
Hardware context matters. Figure 03 incorporates redesigned vision hardware with twice the frame rate, one-quarter the latency, and 60% wider field of view per camera compared with prior generations, plus embedded tactile sensing and palm cameras. These sensors feed the Helix controller directly. The robot’s reduced mass and volume relative to Figure 02 improve maneuverability in typical residential spaces, while soft textiles and multi-density foam address safety in human environments.
Cost, Reliability, and Fleet Economics
Halving task-specific data requirements while expanding coverage thirtyfold changes the unit economics of behavior development. If a service operator previously budgeted for dedicated data collection teams per customer site, the new approach concentrates that spend on a smaller set of representative environments. The $3.5 billion compute commitment signals continued investment in the pretraining stage, where marginal gains in Index scale appear to compound downstream reliability.
Second-order effects include maintenance and update logistics. A shared foundation model allows over-the-air policy updates across a fleet without recollecting data in every home. However, the 40-67% per-task spread indicates uneven reliability that could require fallback behaviors or human supervision in early deployments. Operators will need robust monitoring of failure modes—such as grasp slippage on irregular objects or navigation in cluttered rooms—to maintain service-level agreements.
Counter-arguments note that the 30 homes, while unseen, were all in the same geographic region and may share statistical similarities in layout, lighting, and object types with Bay Area trai
Sources
- https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization
- https://theaiinsider.tech/2026/09/17/figure-unveils-helix-2-5-with-zero-shot-humanoid-generalization-across-30-homes/
- https://www.humanoidsdaily.com/news/figure-helix-2-5-30-unseen-homes
- https://www.unite.ai/figure-introduces-helix-2-5-tested-zero-shot-in-30-unseen-homes/
Who services this hardware
Topics
Manufacturers
Related articles
- Skild AI S1 Drives $100M Revenue Run Rate Ten Months In — Skild AI reached a $100 million annual recurring revenue run rate just ten months after its first commercial deployments of the S1 robotic foundation model. The
- Hugging Face LeRobot GR00T N1.7 Integration and Microduck Biped Launch — Hugging Face updated its LeRobot library in July 2026 to support NVIDIA's GR00T N1.7 VLA model for fine-tuning. In late August, its Pollen Robotics team opened
- Toyota Research Institute Scales LBMs for Manufacturing at Automate 2026 — Toyota Research Institute presented updates on its Large Behavior Models at Automate 2026. Director Erin McColl described diffusion-based multitask policies tha
- Skild AI S1 Achieves 10-Minute Tasks from One Video Prompt — Skild AI released its S1 robotics foundation model on August 25, 2026. The model performs unseen long-horizon manipulation tasks up to 10 minutes long after wat