PASSAGE: Scaling Scene-Aligned Motion Learning for Perceptive Humanoid Traversal in Cluttered Environments
PASSAGE is a perception-conditioned planner-tracker framework for humanoid traversal of cluttered environments. The authors collect 100 hours of scene-aligned human motion with virtual reality and inertial motion capture across 1,500 cluttered scenes; a conditional flow-matching planner generates short-horizon references from motion history, a local destination and a robot-centric multi-layer elevation map, and a perceptive whole-body tracker executes them at 50 Hz with geometric feedback. Real-time chunking promotes inter-chunk consistency, and planner-side reinforcement learning on top of the frozen tracker further improves closed-loop performance. The system needs no skill annotations and no obstacle-specific policies: one planner-tracker pair learns to select and compose stepping, squeezing and ducking on geometries it has never seen. Scaling captured data from 6 to 100 hours raises mean contact-free success on held-out scenes from 48.1% to 68.9%, and 70.3% with validated scene augmentation. The fully onboard system integrates egocentric 3D LiDAR, online occupancy mapping, 6.25 Hz planning and 50 Hz control on a Jetson AGX Orin, traversing 50 unseen physical layouts without prebuilt maps or offboard computation.
Traversal is the precondition for humanoids to leave the lab, and PASSAGE moves the scaling bottleneck away from writing task-specific reinforcement-learning objectives or curating motion libraries and onto collecting more scene-aligned human motion, which is a data-side path that can actually scale. One hundred hours of motion across 1,500 scenes, an explicit 6-to-100-hour scaling curve, and a complete stack that runs entirely onboard on a Jetson AGX Orin and was validated across 50 unseen physical layouts put it clearly ahead of the rest of the batch on completeness, empirical grounding and industrial relevance.