Figure releases Helix 2.5, boosting robot task success rate to 56%

Figure (Figure Technologies) official logo

Figure releases Helix 2.5, boosting robot task success rate to 56%

Figure AI's latest humanoid control model completed household chores in 30 unfamiliar homes without any additional training

A humanoid robot that can walk into a stranger’s house and make the bed without being trained on that specific house sounds like science fiction. Figure AI would like to suggest otherwise.

On September 17, 2026, Figure unveiled Helix 2.5, a new AI control model for its humanoid robots that achieved a 56% success rate across three household tasks performed in 30 previously unseen Bay Area homes. That number comes from 237 successful completions out of 420 total attempts, and it lands against a baseline of 9% for a comparable model trained from scratch on the same tasks.

What Helix 2.5 actually did

The model was tested on three full-body chores: making beds, folding towels, and tidying living rooms. Each task requires the robot to reason about objects it has never encountered before, in spatial configurations it has never seen.

Bed making led the performance numbers, with Figure’s robot succeeding in 94 of 140 attempts, a 67% success rate. Towel folding came in second at 62%, clearing 87 of 140 trials. Tidying toys was the hardest, landing at 40% with 56 successful runs out of 140.

Advertisement

The hardware running all of this is Figure’s 03 platform, the company’s current generation of humanoid robot. Helix 2.5 is the software layer sitting on top of it, handling perception, decision-making, and physical control simultaneously.

The critical qualifier here is zero-shot generalization. The robot received no additional training specific to these homes or the objects inside them.

The dataset behind the jump

The performance leap from 9% to 56% traces back to something Figure calls Index, a pretraining dataset built from recordings of human behavior. Figure publicly launched the Index data platform on August 25, 2026, roughly three weeks before the Helix 2.5 announcement.

Helix 2.5 also builds on a prior iteration. Figure launched Helix 02 in January 2026, establishing the control architecture that Helix 2.5 now extends.

CEO Brett Adcock described Helix 2.5 as the most important project Figure has undertaken.

Why 56% is both impressive and incomplete

A 56% success rate is genuinely remarkable for zero-shot generalization across real, messy, varied household environments. It also means the robot fails almost half the time, which is a long way from the reliability standard that everyday home use would require.

The comparison isn’t a robot versus a human housekeeper. It’s a robot trained on general human behavior data versus the same hardware with no meaningful pretraining. Going from one-in-eleven to better-than-one-in-two, without seeing the environment beforehand, is a significant engineering result.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.
Figure releases Helix 2.5, boosting robot task success rate to 56%
Figure releases Helix 2.5, boosting robot task success rate to 56%

Figure AI's latest humanoid control model completed household chores in 30 unfamiliar homes without any additional training

Figure (Figure Technologies) official logo

A humanoid robot that can walk into a stranger’s house and make the bed without being trained on that specific house sounds like science fiction. Figure AI would like to suggest otherwise.

On September 17, 2026, Figure unveiled Helix 2.5, a new AI control model for its humanoid robots that achieved a 56% success rate across three household tasks performed in 30 previously unseen Bay Area homes. That number comes from 237 successful completions out of 420 total attempts, and it lands against a baseline of 9% for a comparable model trained from scratch on the same tasks.

What Helix 2.5 actually did

The model was tested on three full-body chores: making beds, folding towels, and tidying living rooms. Each task requires the robot to reason about objects it has never encountered before, in spatial configurations it has never seen.

Bed making led the performance numbers, with Figure’s robot succeeding in 94 of 140 attempts, a 67% success rate. Towel folding came in second at 62%, clearing 87 of 140 trials. Tidying toys was the hardest, landing at 40% with 56 successful runs out of 140.

Advertisement

The hardware running all of this is Figure’s 03 platform, the company’s current generation of humanoid robot. Helix 2.5 is the software layer sitting on top of it, handling perception, decision-making, and physical control simultaneously.

The critical qualifier here is zero-shot generalization. The robot received no additional training specific to these homes or the objects inside them.

The dataset behind the jump

The performance leap from 9% to 56% traces back to something Figure calls Index, a pretraining dataset built from recordings of human behavior. Figure publicly launched the Index data platform on August 25, 2026, roughly three weeks before the Helix 2.5 announcement.

Helix 2.5 also builds on a prior iteration. Figure launched Helix 02 in January 2026, establishing the control architecture that Helix 2.5 now extends.

CEO Brett Adcock described Helix 2.5 as the most important project Figure has undertaken.

Why 56% is both impressive and incomplete

A 56% success rate is genuinely remarkable for zero-shot generalization across real, messy, varied household environments. It also means the robot fails almost half the time, which is a long way from the reliability standard that everyday home use would require.

The comparison isn’t a robot versus a human housekeeper. It’s a robot trained on general human behavior data versus the same hardware with no meaningful pretraining. Going from one-in-eleven to better-than-one-in-two, without seeing the environment beforehand, is a significant engineering result.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.