ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation

· Editorial Team estimated
vision-language-action cross-embodiment zero-shot-transfer benchmark manipulation

ZETA is the first controlled study of zero-shot cross-embodiment transfer for Vision-Language-Action models. It separates strict zero-shot transfer (target embodiment absent from all training data) from pretrain-exposed zero-shot transfer (target embodiment seen only in pretraining), and introduces a benchmark spanning 14 held-out target embodiments in simulation and the real world. Controlled factor analysis shows local end-effector representations (+15pp), source embodiment diversity (+18pp), and auxiliary co-training (+7pp) all help, while adding just 5% target-embodiment data during pretraining improves target-embodiment progress by 13.4 percentage points.

Paper · arXiv:2609.02546

As robot hardware evolves and task-specific data collection stays expensive, zero-shot generalization to unseen embodiments is essential for generalizable Vision-Language-Action (VLA) models. Yet the literature lacks a unified definition of zero-shot transfer and controlled evaluation settings that isolate embodiment changes from differences in tasks, scenes, or protocols — making results hard to compare. ZETA is a systematic attempt to close that gap.

Core Innovation

  • Two regimes made explicit — strict zero-shot transfer (the target embodiment is absent from all training data) is distinguished from pretrain-exposed zero-shot transfer (it appears only during pretraining), and the paper argues the two must be reported separately.
  • Controlled benchmark — 14 held-out target embodiments spanning simulation and real-world validation, isolating embodiment change from task, scene, and protocol confounds.
  • Four-factor controlled analysis — state-action representations, pretraining embodiment diversity, auxiliary co-training objectives, and target-embodiment exposure are each varied in a controlled way.

Results

  • Local end-effector (EEF) state-action representations, source embodiment diversity, and auxiliary co-training improve cross-embodiment transfer by roughly 15, 18, and 7 percentage points, respectively.
  • Adding only 5% target-embodiment data during pretraining improves average target-embodiment progress by 13.4 percentage points — direct evidence that strict and pretrain-exposed zero-shot transfer behave differently and should not be conflated.

Limitations

The controlled setting is stationary tabletop manipulation with two-finger grippers; mobile-base control, dexterous hands, and long-horizon tasks are explicitly left to future work, so the measured factor effects should not be extrapolated blindly. Real-world validation is present but limited in scale, and conclusions are drawn within the benchmark’s own protocol.

Industry Implications

The findings translate directly into data strategy for teams building generalist VLA policies: curate embodiment diversity in pretraining, prefer local EEF action representations, and treat a few percent of target-platform data in pretraining as a high-leverage investment when onboarding a new robot. Equally important is evaluation discipline — vendors should not advertise pretrain-exposed results as strict zero-shot transfer. A shared protocol of this kind gives buyers a fairer basis for comparing commercial robot foundation models.