FolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects
FolDeX is a physical-world benchmark built entirely from real-robot data, centered on garment folding, that targets long-horizon deformable-object manipulation where policies must track changing states and execute reliable multi-stage bimanual interactions. Because real-robot data is costly, it studies efficient reuse of heterogeneous physical experience along four axes: human intervention and recovery data collected during deployment; transfer across tasks including garment categories and rigid-to-deformable manipulation; reuse across scenes with changing lighting, background, and layout; and transfer across embodiments. The benchmark provides 2,000+ hours of real-robot data spanning 20+ tasks and 10+ embodiments, plus a fair evaluation platform for externally submitted policies with standardized tasks, held-out physical objects, controlled initializations, and a unified execution protocol.
Paper · arXiv:2609.10243Embodied AI has an evaluation problem. Vision-language-action and world-action models post strong numbers in simulation, yet degrade sharply on real robots — nowhere more so than in long-horizon deformable-object manipulation, where a policy must hold a changing fabric state in mind across many stages of bimanual interaction. The existing real-robot benchmarks that could catch this failure mode have focused largely on short-horizon rigid-object tasks. Deformable manipulation, and the data economics behind it, has been under-benchmarked.
Core Innovation
FolDeX is a physical-world benchmark built entirely from real-robot data, with garment folding as its primary task. Rather than adding another simulator, it confronts the central practical constraint: real-robot data is expensive, so the field needs to understand how to reuse heterogeneous physical experience efficiently. The benchmark is organized around four research axes:
- Human intervention and recovery data collected during deployment — turning the messy, corrective moments of real operation into supervision.
- Cross-task transfer, including between garment categories and from rigid- to deformable-object manipulation.
- Cross-scene reuse, under changes in lighting, background, and layout.
- Cross-embodiment transfer.
Together these axes make the dataset a laboratory for data efficiency, not just a leaderboard.
Results
- 2,000+ hours of real-robot data.
- 20+ tasks and 10+ embodiments.
- A fair real-robot evaluation platform for externally submitted policies: standardized tasks, held-out physical objects, controlled initializations, and a unified execution protocol.
- The platform is publicly accessible at https://ai.midea.com/#/fold-challenge.
Limitations
The abstract describes the benchmark’s construction and research axes but reports no quantitative baselines or success rates, so its discriminative power across policies is not yet demonstrated in the summary. Coverage is concentrated on garment folding, and while rigid-to-deformable transfer is an axis, the benchmark’s breadth beyond clothing is unclear. As with any hosted challenge, reproducibility depends on the continued availability of the physical platform and held-out objects.
Industry Implications
Data, not model architecture, is increasingly the bottleneck for generalist robot policies. A curated, multi-embodiment, multi-task corpus of 2,000+ hours of real folding data — with a standardized, held-out evaluation protocol attached — is exactly the kind of shared infrastructure that lets the field compare policies honestly and lets labs with limited robot fleets piggyback on someone else’s physical experience. For companies building manipulation policies for laundry, textiles, packing, or any deformable workflow, FolDeX sketches the data-reuse playbook they will need.