Figure AI Unveils Index Model and Project Go-Big for Internet-Scale Humanoid Pretraining
Robot Details
Robotics & AI News • OriginOfBotsPublished
August 25, 2026
Reading Time
2 min read
Author
Origin Of Bots Editorial Team

Announcement of the Index Humanoid Foundation Model
Figure AI officially announced Index, a massive multimodal foundation model tailored specifically for end-to-end humanoid perception and motor control. Index bridges high-level semantic reasoning with low-level torque commands, enabling robots to interpret natural instructions and translate them into physical motions.
Project Go-Big: Internet-Scale Pretraining Strategy
Central to the announcement is Project Go-Big, an initiative aimed at ingesting web-scale human video data and translating observed human biomechanics directly into robot policy training. This approach overcomes the bottleneck of manual teleoperation by using passive video demonstrations to seed physical intuition.
Direct Human-to-Robot Motion Mapping
Through novel vision-based retargeting algorithms, Figure demonstrated how human manipulation behaviors—such as opening latches, packing boxes, and handling delicate tools—can be transferred directly onto humanoid kinematic trees without requiring manual kinematic joint remapping.
16-Degree-of-Freedom Dexterous Manipulation
Figure’s latest hand designs feature integrated tactile sensors across every fingertip and 16 degrees of freedom. This allows Figure humanoids to grasp objects with variable compliance, modulating grip firmness dynamically based on slip feedback and surface texture.
Field Validation in Automotive Manufacturing Lines
Building upon its trial deployments at the BMW Spartanburg manufacturing facility, Figure shared reliability metrics demonstrating that humanoid units achieved high uptime placing sheet metal components and sorting parts with millimeter-level positional repeatability.
Continuous Multi-Hour Autonomous Shift Demonstrations
Figure showcased extended multi-hour continuous autonomous shifts where humanoids managed parts sorting, self-navigated to charging cradles, and resumed duties without human teleoperation intervention, establishing a benchmark for operational reliability.
Vision-Language-Action (VLA) Architecture Convergence
The architecture consolidates vision processing, language comprehension, and motor planning into a unified transformer backbone. This prevents communication latency between isolated perception and actuation stacks, enabling fluid reaction times to human co-workers.
Scaled Production Targets and Hardware Maturation
Figure emphasized its progression toward high-rate assembly line production. By optimizing component count and utilizing automotive-grade manufacturing practices, the company aims to significantly reduce the per-unit bill of materials over the coming production cycles.
Sources
Learn More About This Robot
Discover detailed specifications, reviews, and comparisons for Robotics & AI News.
View Robot Details →