Gasgoo Munich- Maniformer's 20,000th MEgo disembodied data collection device officially rolled off the production line on August 31— marking the industry's first instance of mass production for such equipment. Simultaneously, Maniformer announced it has generated over 1 million hours of high-quality disembodied data, becoming the first data provider in the industry to offer a million-hour dataset available for commercial sale.
At the ceremony, Maniformer Chairman and CEO Yao Maoqing handed over the 20,000th MEgo unit to JD.com. The device is set to enter real-world operational scenarios immediately.
Separately, Maniformer announced a partnership with Tencent's Robotics X Lab to provide data support for the development of Tencent's embodied models.
Maniformer's MEgo device introduces an industry-first multi-view, multi-modal approach. With a panoramic view exceeding 300 degrees and wrist close-ups, the system ensures robots see both the big picture and fine details like fingertips. Spatially, 3D tactile sensing paired with millimeter-level trajectory reconstruction allows every grasp to be precisely replicated. Temporally, hardware-level synchronization across all channels and sub-millisecond alignment perfectly sync vision, touch, and posture, eliminating "spatiotemporal misalignment."
The MEgo also features a wireless, lightweight design, allowing it to follow humans into every corner of the real world.

Image Source: Maniformer
The 1 million hours of disembodied data accumulated by Maniformer includes multi-modal information such as RGB, Depth, IMU, tactile, and audio data. This forms a scaled data supply covering real-world scenarios, actions, and interactions.
Regarding data collection formats, Maniformer has established three data systems: Ego Bare Hand, Ego+Wrist, and Ego+UMI Gripper. These correspond to natural human operation, wrist posture and continuous motion trajectories, and robot end-effector execution, respectively. Specifically, Ego Bare Hand data totals approximately 500,000 hours, Ego+Wrist around 350,000 hours, and Ego+UMI Gripper about 150,000 hours.
All data is sourced from real-world open environments, covering 22 major scenario categories, over 10,000 real environments, more than 50,000 object types, and over 500 specific tasks. It spans eight high-frequency scenarios—including residential, office, retail, manufacturing, warehousing, dining, entertainment, and transportation—and extends further into long-tail scenarios such as public services, logistics, healthcare, education, elderly care, cultural tourism, automotive services, and agricultural production.

Image Source: Maniformer
The eight top-tier scenarios contribute approximately 82% of the data, forming the core foundation for general physical AI capabilities. The remaining 14 long-tail scenarios further expand data coverage across vertical industries.
Behind Maniformer's million-hour dataset lies a full-link technology system supporting the continuous production of real-world data.
To accurately reconstruct human hand movements and capture full-body posture, Maniformer built the MEgo Engine data processing infrastructure. Covering data processing, perception, annotation, and quality assessment, it uses high-precision pose reconstruction and motion mapping to convert real-world information on hands, bodies, heads, robot end-effectors, and the environment into unified spatial data.
At the event, Maniformer further released HandPose, a high-precision hand reconstruction technology. Utilizing five-eye panoramic perception, robust hand tracking, three-eye deep fusion, parametric hand reconstruction, and high-precision trajectory recovery, the technology achieves robust hand reconstruction across all scenarios.
Beyond hands, Maniformer has also developed spatial reconstruction capabilities for the head, body, robot end-effectors, and the environment.
To ensure efficient utilization of massive amounts of real-world data, Maniformer has established the ManiEval multi-dimensional data quality inspection and grading system.
ManiEval evaluates data across multiple dimensions—including data type, task type, sensor quality, semantic quality, and action quality. Adopting a "grade, don't filter" approach, it matches data of varying quality and characteristics to different model training needs, ensuring every piece of real data finds its suitable application scenario.

Image Source: Maniformer
Building on this, Maniformer continues to explore applications for different data types in embodied brains, Vision-Language-Action (VLA) models, and World Action Models. This transforms real-world data into machine perception, action, and reasoning capabilities, thereby bridging the complete loop of "collection—reconstruction—quality inspection—training—deployment—feedback."
Currently, Maniformer's million-hour disembodied dataset is serving top-tier clients such as Tencent's Robotics X Lab and Robbyant, providing tailored data products and services based on specific model training requirements.









