Gasgoo Munich- At the 2026 World Artificial Intelligence Conference (WAIC 2026), DexForce officially released and open-sourced Aligned DexWorld. It marks the world's first large-scale dataset bridging human demonstrations, simulation, and real-world robotics, achieving three-dimensional alignment across time, space, and action. By providing standardized, reusable, and composable data infrastructure, the release aims to help embodied AI models break through scaling bottlenecks.
Designed as a high-quality heterogeneous dataset for the large-scale pre-training of embodied models, Aligned DexWorld integrates mainstream public datasets—including Droid, AgiBotWorld-Beta, RoboMIND, RH20T, LIBERO, RoboTwin, and VITRA—into a unified, standardized post-processing pipeline. The collection spans diverse data types, ranging from human demonstrations and real-world teleoperation to synthetic simulation and first-person perspectives.
In terms of scale, Aligned DexWorld comprises more than 2 million human and robot data samples. These include real-world data from various hardware forms—such as humanoid robots and robotic arms—alongside simulation data and first-person hand demonstrations.

Image Source: DexForce
Unlike traditional datasets that merely standardize formats, Aligned DexWorld’s core value lies in deep, physics-level alignment. By achieving unified physical alignment across space, action, and time for all heterogeneous data, it ensures that the 3D world observed by the model corresponds perfectly with the robot's executed actions.
To handle the alignment requirements of 2 million heterogeneous data points, the dataset employs a closed-loop standardization workflow involving manual calibration, geometric solving, AI-assisted processing, and human verification. By combining visual calibration tools with algorithm optimization, the system precisely corrects data deviations and avoids the pitfalls of fully automated processing. This ensures the alignment logic for every data point remains traceable, verifiable, and reproducible.
Recognizing the prohibitive cost of manually segmenting and labeling massive datasets, DexVerse developed a fully automated annotation pipeline divided into two stages: behavioral segmentation and semantic labeling. First, candidate segments are generated based on robot movement speed and gripper status. A Vision-Language Model (VLM) then merges these candidates, defines standard subtasks, and automatically generates text annotations—achieving end-to-end automation.
This automated labeling approach drastically reduces labor costs while efficiently churning out massive volumes of standardized, aligned subtask data. By providing high-quality, modular data, the solution significantly enhances the generalization capabilities of embodied models and reduces the data required for training new tasks. Ultimately, it lays a low-cost, reusable foundation for data production across the robotics industry, accelerating the large-scale deployment of general physical AI.








