Brain · AI-derived

Figure AI Helix 2.5 Delivers 56% Zero-Shot Success Across 30 Homes

Figure AI released Helix 2.5 on September 17, 2026. The Index-pretrained model enabled Figure 03 humanoids to tidy rooms, fold towels, and make beds in 30 unseen Bay Area homes with no new data or fine-tuning, raising full-task success from 9% to 56%.

Figure AI Helix 2.5 Delivers 56% Zero-Shot Success Across 30 Homes

ZeroGantry analysis

At 56% full-task success after halving adaptation data and scaling to 30 homes, Helix 2.5 implies roughly 2x lower behavior development cost per deployed unit once fleets exceed a few dozen robots. Operators should watch continued Index scaling for reliability gains above 80% before committing to unsupervised home pilots; ignore only if per-home data collection remains cheaper than the $3.5B pretraining path for their use case.

Helix 2.5 and the Zero-Shot Generalization Milestone

On September 17, 2026, Figure AI announced Helix 2.5, a vision-language-action foundation model pretrained on its Index human-behavior dataset. The release detailed results from 420 trials in 30 rented Bay Area homes never encountered during training or adaptation. Three whole-body tasks—tidying living rooms by placing 13-15 toys into a basket, folding four towels into a basket, and making beds by positioning pillows and smoothing comforters—were executed using a single fixed checkpoint per behavior. No data collection, weight updates, or environment-specific fine-tuning occurred in the test homes.

The controlled comparison isolated Index pretraining as the variable. A policy trained from scratch achieved 9% full-task success under identical task data, architecture, and evaluation criteria. The Index-pretrained Helix 2.5 reached 56%, completing 237 of 420 trials with no partial credit. Per-task rates showed variation: bed making at 67% (94/140), towel folding at 62% (87/140), and tidying at 40% (56/140). This gap directly attributes the generalization lift to pretraining on diverse human video rather than task-specific data alone.

Figure also compared Helix 2.5 against a prior Helix 02 policy trained with data collected directly in its evaluation environment. Helix 2.5 matched that policy’s success rate while requiring only half the task-specific adaptation data and generalizing across 30 homes instead of one. Behavior specification costs dropped by a factor of two while the deployment scope expanded thirtyfold. These efficiency gains address a central economic barrier for household humanoid deployment, where per-home data collection has historically scaled poorly.

Index Dataset Pretraining and Scaling Laws

Index, introduced publicly on August 25, 2026, aggregates human behavior video captured via a mobile app network. At launch it reported 264,000 downloads across 108 countries, over 44,000 weekly active users, 16 million videos uploaded, and $15 million paid to contributors. Processing reached 30 minutes of video per second, equivalent to roughly 4.9 years of human activity daily. Per 1,000 hours collected, the dataset contained 373 unique tasks, 1,146 manipulated objects, and 116 environments after a five-stage pipeline of filtering, fraud review, deduplication, rebalancing, and annotation.

Helix 2.5 pretraining on Index produced a measurable human-to-robot transfer scaling law. Doubling pretraining data volume improved downstream robot-action prediction loss predictably enough for the team to forecast the largest training run’s final loss to four decimal places before execution. At the time of the Helix 2.5 announcement, Index generated approximately 35 minutes of new human experience data every second. Figure committed $3.5 billion in compute through Nscale, targeting up to 100,000 NVIDIA Vera Rubin GPUs for continued scaling.

No single evaluation task exceeded 1.90% of the Index pretraining distribution, ensuring broad coverage rather than narrow overfitting. This diversity supports the observed transfer: policies adapted from the same base model handled locomotion, rigid and deformable manipulation, bimanual coordination, and active perception across novel layouts and objects.

Implications for Household Humanoid Deployment

Current commercial and research humanoids typically require environment-specific data collection and fine-tuning before reliable operation in new settings. Helix 2.5’s results indicate that large-scale human-behavior pretraining can reduce or eliminate that step for a defined set of chores. Service providers could deploy fleets with a common base model, then adapt behaviors once using modest task data collected elsewhere, cutting per-site onboarding costs and time.

The 56% full-task success rate remains well below levels needed for unsupervised home service. Strict criteria—no partial credit, safety interventions counted as failures, and timeouts—highlight remaining failure modes in perception, grasp stability, and long-horizon planning under novel lighting, clutter, or object configurations. Bed making outperformed toy tidying, suggesting deformable object handling and precise spatial reasoning still vary by task complexity.

Hardware context matters. Figure 03 incorporates redesigned vision hardware with twice the frame rate, one-quarter the latency, and 60% wider field of view per camera compared with prior generations, plus embedded tactile sensing and palm cameras. These sensors feed the Helix controller directly. The robot’s reduced mass and volume relative to Figure 02 improve maneuverability in typical residential spaces, while soft textiles and multi-density foam address safety in human environments.

Cost, Reliability, and Fleet Economics

Halving task-specific data requirements while expanding coverage thirtyfold changes the unit economics of behavior development. If a service operator previously budgeted for dedicated data collection teams per customer site, the new approach concentrates that spend on a smaller set of representative environments. The $3.5 billion compute commitment signals continued investment in the pretraining stage, where marginal gains in Index scale appear to compound downstream reliability.

Second-order effects include maintenance and update logistics. A shared foundation model allows over-the-air policy updates across a fleet without recollecting data in every home. However, the 40-67% per-task spread indicates uneven reliability that could require fallback behaviors or human supervision in early deployments. Operators will need robust monitoring of failure modes—such as grasp slippage on irregular objects or navigation in cluttered rooms—to maintain service-level agreements.

Counter-arguments note that the 30 homes, while unseen, were all in the same geographic region and may share statistical similarities in layout, lighting, and object types with Bay Area trai

Sources

Who services this hardware

Topics

Manufacturers

Related articles

Editorial methodology