HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
HumanTracker is a large-scale benchmark containing approximately 153 hours of professional optical motion trajectories for evaluating humanoid motion tracking, paired with HumanScore, a preference-aligned metric trained on 12K motion pairs that better predicts human perception and reveals contact and stability failures that kinematic metrics miss.
Paper · arXiv:2608.13555Humanoid motion tracking is the backbone of teleoperation and whole-body imitation learning, yet the way we evaluate it often disagrees with what people actually see in videos. Classic kinematic error metrics average per-frame pose differences and completely miss the physical artifacts that matter most in practice: unstable support, foot skating, and mistimed touch-downs. At the same time, the test suites researchers actually use are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. The result is a measurement problem — models can score well on metrics while looking wrong to a human observer.
Core Innovation
HumanTracker attacks both halves of this problem. First, it is a benchmark built on roughly 153 hours of optical motion trajectories recorded from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. Second, the authors introduce HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions, which learns what human viewers consider good tracking rather than relying on hand-designed geometric error.
The key design bet is that evaluation should be perceptually aligned: instead of asking “how far is the estimated pose from ground truth per frame,” it asks “would a person find this tracking acceptable?” This makes the metric sensitive to contact and stability failures — skating feet, late touch-downs, wobbly support — that kinematic averages wash out.
Results
- Benchmark scale: approximately 153 hours of optical motion data, four motion families, text-labeled for fine-grained diagnosis
- HumanScore trained on 12K motion pairs (24K motions)
- Across representative state-of-the-art trackers, HumanScore better predicts human preferences than kinematic metrics
- The metric reveals contact and stability failures that conventional kinematic measures miss
Limitations
The abstract describes the benchmark and metric’s design and validation but does not publish per-tracker scores or detailed benchmark results in the abstract itself; those will live in the full paper. The metric’s preference alignment is learned from a specific annotation pool, so its generality to other motion families, performer styles, and evaluation contexts still needs scrutiny. HumanScore measures tracking quality as perceived by humans — it complements, but does not replace, task-level success metrics for downstream imitation or teleoperation performance.
Industry Implications
For the humanoid industry, this is a measurement-infrastructure paper with real leverage: teleoperation and whole-body imitation pipelines are only as good as the evaluation that drives their iteration. A scalable, human-aligned tracking metric gives developers a tool to catch skating feet and unstable contact during data collection and policy evaluation — precisely the failures that look fine in an error plot but break physical robots. As humanoid platforms move toward long-horizon, contact-rich behaviors, benchmarks like HumanTracker provide the standardized, perceptually grounded yardstick the field has been missing.