HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

· Editorial Team estimated
robot learning imitation learning data collection UMI manipulation

HiFi-UMI raises the fidelity of robot-free UMI data collection to 3 mm end-effector accuracy using head-mounted stereo-inertial SLAM, native relative pose estimation, and wide-angle stereo cameras. Policies post-trained solely on this data match in-domain teleoperation across three backbone families, and a 4000-hour pretraining corpus lowers action error on unseen tasks by 41%.

Paper · arXiv:2607.25895

Background

Deployable robot manipulation policies are bottlenecked by the scarcity of high-quality training data. Real-robot teleoperation is accurate but slow and expensive to scale. Robot-free data collection via UMI (Universal Manipulation Interface) scales readily, but current practice treats it only as pretraining material, requiring a small real-robot “anchor” for post-training. This leaves the central question open: can robot-free data alone produce deployable policies?

Core Innovation

HiFi-UMI is a portable data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view. Key design choices include: head-mounted offline stereo-inertial SLAM for accurate camera egomotion, native rather than reconstructed relative gripper pose estimation, a shared microsecond GPIO trigger across all cameras, and two wide-angle cameras per hand covering approximately 200 degrees. The system achieves 3 mm workspace-local end-effector accuracy without any external tracking infrastructure.

The critical result is zero-robot post-training: a policy fine-tuned solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbone families spanning vision-language-action (StarVLA-QwenPI, OpenPI-pi_0.5) and world-action-model (LingBot-VA) architectures.

Results

Success-rate differences between HiFi-UMI-trained and teleoperation-trained policies are -2.5, +3.1, and -0.6 percentage points respectively. The strongest policy reaches 85% on a precision insertion task, despite no HiFi-UMI trajectory being collected in the evaluation scene while the teleoperation baseline is. Pretraining on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by an additional 18.1 percentage points. The authors open-source HiFi-UMI-2K — 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations.

Limitations

The system requires head-mounted hardware (SLAM camera rig) and offline post-processing for trajectory reconstruction, limiting real-time streaming applications. The current validation is on pick-and-place and insertion tasks; generalization to more diverse manipulation scenarios (dexterous, contact-rich) remains to be demonstrated.

Industry Implications

For robotics teams, HiFi-UMI’s zero-robot post-training result is transformative: it implies that large-scale, high-quality robot-free data collection can substitute for expensive and time-consuming real-robot teleoperation for policy training. The open-source release of 2,000 hours of validated demonstrations provides a substantial new resource for the community. The 41% action error reduction on unseen tasks from pretraining further suggests a viable path toward generalist manipulation models trained primarily from robot-free data.