AdaRoboVLG: Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis
AdaRoboVLG is a task-adaptive Vision-Language-Grasp (VLG) framework supporting generalizable grasp synthesis across different robotic hands. Unlike VLG methods that tightly couple foundation models with end-to-end grasp policies, it learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while task-dependent understanding is offloaded to specialized foundation-model modules that provide composable spatial, cognitive, and temporal priors — enabling contextually adaptive grasping without retraining the underlying grasp policy. Extensive simulation and real-world experiments show efficient base-policy learning with strong cross-hand generalization, effective use of the three priors on representative grasping challenges without compromising synthesis performance versus state-of-the-art methods, and joint operation of the priors for functional grasping in cluttered and dynamic environments.
Paper · arXiv:2609.04096Existing vision-language grasping systems weld foundation models onto end-to-end grasp policies, so task understanding and physical grasp synthesis advance or stall together. AdaRoboVLG separates the two: a stable, generalizable grasp-synthesis core that can be reused across hands, plus pluggable foundation-model priors that make each grasp contextually appropriate.
Core Innovation
- Decoupled architecture — an efficient base policy generates and evaluates physically feasible grasp candidates via explicit kinematic mapping and force-closure-based stability estimation, independent of task semantics.
- Composable foundation priors — spatial, cognitive, and temporal priors from foundation models are integrated into grasp synthesis, enabling task-adaptive behavior without retraining the grasp policy.
- Scalable paradigm — future foundation-model improvements can be plugged in as new priors, directly upgrading grasp capability without policy redesign.
Results
- Efficient learning and strong cross-hand generalization of the base policy in simulation and real-world experiments.
- Spatial, cognitive, and temporal priors each address a representative grasping challenge without compromising grasp synthesis performance compared to state-of-the-art methods.
- Priors operate jointly to achieve functional grasping in cluttered and dynamic environments.
Limitations
The abstract summarizes benchmark and real-world demonstrations but does not quantify success rates or the number of hands and objects tested; the composable-prior mechanism depends on the quality and latency of the underlying foundation models; and long-horizon or whole-body grasping scenarios beyond the reported settings remain open.
Industry Implications
Grasping remains the highest-failure-point operation in warehouse, logistics, and service robotics. AdaRoboVLG’s decoupling means a robot fleet with different grippers or hands can share one grasp-synthesis core while task awareness is upgraded independently — cutting the cost of adapting to new hardware and new tasks, and giving integrators a concrete route to keep pace with fast-moving foundation models without redeveloping grasp stacks.