MATE: Multi-Agent Virtual Teleoperation Platform for Humanoid Collaboration Data Collection
Humanoid robots need embodied experience of collaborative work, but physical multi-robot data collection does not scale: it demands several robots, a dedicated space, and repeated resets. MATE replaces the hardware with a shared physics simulation in which multiple geographically distributed operators simultaneously teleoperate whole-body humanoids, preserving the physically coupled interactions between humanoids, objects, and environments that make collaboration data valuable in the first place. Using MATE the authors build a multi-humanoid collaboration dataset of 24.1 hours across 2,500 joint episodes and five long-horizon tasks, including object handover, relay delivery, environment interaction, and cooperative transport, and introduce EAIS, an execution-aligned interaction sampling strategy that prioritizes task-progressing and interaction-critical behavior. Imitation learning and vision-language-action policies trained on this data learn effectively and transfer zero-shot to a physical humanoid without real-world fine-tuning.
Paper · arXiv:2609.26520Humanoid loco-manipulation and collaboration skills are learned from embodied experience, and collaboration is the part that does not scale. Collecting data for a single humanoid already means hardware, space, and resets; collecting data for two or three humanoids interacting with the same object means all of that multiplied, plus scheduling every operator and every robot into one room at one time. The result is that multi-humanoid data stays scarce while the tasks that most need it — handing an object to a partner, carrying something jointly, reacting to another agent’s body — are precisely the ones a single-agent pipeline cannot synthesize. MATE takes the position that the interaction must be physical, but the hardware need not be.
Core Innovation
- Distributed multi-operator teleoperation in one shared physics environment. Multiple operators, in different locations, control whole-body humanoids inside a single physics-based scene. Physical coupling among the humanoids, the manipulated objects, and the environment is preserved, so the collected demonstrations contain genuine contact and coordination rather than scripted playback.
- No robot fleet, no co-location, no repeated physical resets. The expensive parts of multi-robot collection are removed while the interaction structure is retained, which is what makes the platform usable for long-horizon collaborative skills.
- EAIS — Execution-Aligned Interaction Sampling. Within an execution-aligned prefix of an episode, EAIS computes sampling signals and prioritizes task-progressing and interaction-critical behaviours, so training weight concentrates on the moments where coordination actually happens instead of being spread uniformly over long episodes.
- A collaboration dataset built to be learned from. 24.1 hours of coordinated behaviour across 2,500 joint episodes and five long-horizon tasks: object handover, relay delivery, environment interaction, and cooperative transport, among others.
Results
- Scale: a multi-humanoid collaboration dataset of 24.1 hours, 2,500 joint episodes, and five long-horizon tasks, collected through multi-operator virtual teleoperation.
- Policy learning: representative imitation learning and vision-language-action policies trained on this data learn effectively across the diverse collaboration tasks, and EAIS improves learning from interaction-rich demonstrations.
- Transfer: policies exhibit zero-shot transfer from virtual demonstrations to a physical humanoid without real-world fine-tuning, and the authors report efficient data collection as a practical property of the platform.
Limitations
The abstract does not state how many humanoids or operators are involved in a given session, how many physical robots the transfer was validated on, or what the success rate of the transferred behaviour actually is. Zero-shot transfer is claimed without a quantified sim-to-real performance gap, which is the number that would decide how much real data the platform can genuinely replace. Teleoperation quality across distributed operators is also a variable: latency, operator skill, and differences in teleoperation hardware all shape the demonstrations, and none of that is characterised. There is no comparison against data collected on physical multi-robot setups, so the fidelity cost of the virtual medium is unmeasured, and long-horizon tasks are reported by task family rather than by per-task success or recovery behaviour.
Industry Implications
The practical bottleneck for collaborative humanoids is not model capacity but the supply of coordinated multi-robot demonstrations, and every existing route to that supply is capital-intensive: more robots, more floor space, more operator hours per session. A virtual platform that keeps physical coupling while removing the fleet turns data collection from a hardware programme into a software one, which is the kind of shift that lets smaller teams participate. The open question is the sim-to-real accounting: if zero-shot transfer holds at useful success rates, simulated multi-humanoid collection becomes the default acquisition channel for collaboration skills; if it requires a real-data correction pass, the platform is still cheaper but no longer free, and the ratio between the two becomes the deciding operational metric.