One Hand Watches The Other: Dynamic Multi-Agent Cooperation for Sample-Efficient Bimanual Manipulation in Dynamic Environments

· Editorial Team estimated
bimanual-manipulation multi-agent multi-stream-policies dynamic-environments sample-efficiency

DynaMAC resolves the causal limitation of multi-stream policies in dynamic settings by treating the opposite arm as a dynamic task parameter, enabling unified dynamic manipulation and bimanual coordination without an explicit leader-follower relationship. It outperforms leading baselines by 35+ percentage points while requiring 20× fewer samples, and generalizes zero-shot from static demonstrations to dynamic environments.

Paper · arXiv:2607.22119

Background

Multi-stream robot manipulation policies achieve unparalleled sample efficiency and generalization by modeling actions relative to environmental reference frames. However, existing approaches assume these frames are strictly exogenous. This causal assumption collapses in dynamic settings — when a single arm manipulates a moving object or two arms coordinate, each arm becomes part of the other’s dynamic environment. No prior method has provided a unified formulation for dynamic manipulation and bimanual coordination without requiring an explicit leader-follower relationship.

Core Innovation

DynaMAC is a lightweight, policy-agnostic framework that treats the opposite arm as a dynamic task parameter, providing a unified formulation for dynamic manipulation and bimanual coordination without an explicit leader-follower relationship. This resolves the causal limitation of existing multi-stream policies while preserving their sample efficiency, computational speed, and flexibility. The framework is designed as a simple plug-in that can be added to any existing multi-stream policy architecture.

Key Results

Across both dynamic environments and bimanual manipulation tasks on the new DynaBench benchmark, DynaMAC outperforms leading probabilistic and generative baselines by over 35 percentage points while requiring 20× fewer samples. Crucially, DynaMAC generalizes zero-shot from static demonstrations to dynamic environments, substantially simplifying data collection and establishing an elegant bridge toward human-robot collaboration. The DynaBench benchmark itself is a contribution, providing standardized evaluation for dynamic and bimanual manipulation.

Limitations

DynaMAC’s performance depends on accurate state estimation of the dynamic objects or partner arm. While the framework is policy-agnostic, its benefits have been demonstrated primarily on trajectory-level policies; compatibility with end-to-end learned visuomotor policies requires further investigation. The zero-shot generalization from static to dynamic settings, while impressive, was evaluated on a limited set of task configurations.

Industry Implications

Bimanual manipulation is a critical capability for advanced manufacturing, assembly, and logistics — tasks where two arms must coordinate to handle large, flexible, or articulated objects. DynaMAC’s 20× sample efficiency and zero-shot static-to-dynamic generalization mean that roboticists can collect demonstrations in simple static setups and deploy directly in dynamic production environments. This significantly lowers the barrier to deploying bimanual systems in warehouse order fulfillment, automotive assembly, and collaborative human-robot workspaces.