RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

· Editorial Team estimated
embodied foundation model VLA humanoid cross-embodiment manipulation

RynnBrain 1.1 introduces an embodied foundation model family spanning 2B to 122B parameters with contact-point prediction and native 3D grounding. It outperforms all proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench, and demonstrates cross-embodiment VLA deployment on Unitree G1, Astribot-S1, and Tianji-Wuji robots.

Paper · arXiv:2607.17977

Background

Embodied foundation models aim to provide general-purpose perception, spatial reasoning, localization, and planning capabilities for robots. Previous models, however, lacked explicit modeling of contact points critical for fine manipulation, and differences in action spaces across embodiments made it difficult for a single model to transfer directly across robot platforms — limiting both practical utility and generalizability.

Core Innovation

RynnBrain 1.1 introduces two key improvements over version 1.0: contact-point prediction — the model predicts expected contact locations between the end-effector and objects, aligning outputs more directly with manipulation; and native 3D grounding — upgrading spatial reasoning from implicit encoding to explicit 3D coordinate prediction for the 2B and 9B models. These additions significantly improve precision in manipulation tasks.

The newly developed RynnBrain-VLA employs a unified cross-embodiment action space with embodiment-specific masking, enabling direct deployment on Unitree G1 (humanoid), Astribot-S1 (dual-arm desktop robot), and Tianji-Wuji (full-size dual-arm humanoid).

Results

The 122B-A10B variant achieves top performance across VSI-Bench (visual spatial intelligence), MMSI (multimodal spatial understanding), and RefSpatial-Bench (referential spatial grounding), surpassing all evaluated proprietary and open-source models. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs in both success rate and process scores. Joint multi-task and multi-embodiment training further improves performance over per-task training.

Limitations

The paper does not discuss inference latency or computational cost for the 122B model in detail. The scalability of the unified action space to a larger set of embodiments remains unexamined. Contact-point prediction is evaluated primarily for static grasping scenarios; performance under dynamic manipulation is not addressed.

Industry Impact

RynnBrain 1.1 demonstrates that scaling embodied foundation models — through larger capacity, cross-embodiment training, and structured spatial outputs — yields synergistic gains. Its unified multi-embodiment deployment capability provides a viable technical pathway for both industrial and service humanoid robots.