HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

· Editorial Team estimated
humanoid motion-tracking benchmark teleoperation evaluation

HumanTracker is a large-scale benchmark containing approximately 153 hours of professional optical motion trajectories for evaluating humanoid motion tracking, paired with HumanScore, a preference-aligned metric trained on 12K motion pairs that better predicts human perception and reveals contact and stability failures that kinematic metrics miss.

Paper · arXiv:2608.13555

Humanoid motion tracking is the backbone of teleoperation and whole-body imitation learning, yet the way we evaluate it often disagrees with what people actually see in videos. Classic kinematic error metrics average per-frame pose differences and completely miss the physical artifacts that matter most in practice: unstable support, foot skating, and mistimed touch-downs. At the same time, the test suites researchers actually use are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. The result is a measurement problem — models can score well on metrics while looking wrong to a human observer.

Core Innovation

HumanTracker attacks both halves of this problem. First, it is a benchmark built on roughly 153 hours of optical motion trajectories recorded from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. Second, the authors introduce HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions, which learns what human viewers consider good tracking rather than relying on hand-designed geometric error.

The key design bet is that evaluation should be perceptually aligned: instead of asking “how far is the estimated pose from ground truth per frame,” it asks “would a person find this tracking acceptable?” This makes the metric sensitive to contact and stability failures — skating feet, late touch-downs, wobbly support — that kinematic averages wash out.

Results

  • Benchmark scale: approximately 153 hours of optical motion data, four motion families, text-labeled for fine-grained diagnosis
  • HumanScore trained on 12K motion pairs (24K motions)
  • Across representative state-of-the-art trackers, HumanScore better predicts human preferences than kinematic metrics
  • The metric reveals contact and stability failures that conventional kinematic measures miss

Limitations

The abstract describes the benchmark and metric’s design and validation but does not publish per-tracker scores or detailed benchmark results in the abstract itself; those will live in the full paper. The metric’s preference alignment is learned from a specific annotation pool, so its generality to other motion families, performer styles, and evaluation contexts still needs scrutiny. HumanScore measures tracking quality as perceived by humans — it complements, but does not replace, task-level success metrics for downstream imitation or teleoperation performance.

Industry Implications

For the humanoid industry, this is a measurement-infrastructure paper with real leverage: teleoperation and whole-body imitation pipelines are only as good as the evaluation that drives their iteration. A scalable, human-aligned tracking metric gives developers a tool to catch skating feet and unstable contact during data collection and policy evaluation — precisely the failures that look fine in an error plot but break physical robots. As humanoid platforms move toward long-horizon, contact-rich behaviors, benchmarks like HumanTracker provide the standardized, perceptually grounded yardstick the field has been missing.