Gasgoo Munich- July 2026, Shanghai World Expo Exhibition Center. Step into the embodied intelligence zone, and the air is thick with a mix of anxiety and mania. Two years ago at WAIC, only 18 companies dared to test the waters; last year that figure surged to 80; this year, exhibitors have quietly surpassed 200. Yet behind the bustling booths, industry anxiety is mounting: as of early 2026, global high-quality embodied interaction data totals just 500,000 hours, while the entry barrier for training general embodied models sits at the 10 million-hour mark—a gap exceeding 99%.
With data collection solutions proliferating, why does the gap remain so vast? Can the endless stream of technical approaches truly bridge the industry's data chasm? Moving from "a plethora of solutions" to "sufficient data" requires the embodied intelligence sector to clear multiple hurdles.
Over 20 Players Crowd the Field as Four Major Data Collection Paths Emerge
At this year's WAIC embodied intelligence zone, nearly every robotics firm arrived with its own data "production tools." From teleoperation to wearables, and body-free setups to simulation, the collection methods on display cover nearly every path imaginable.

Teleoperation — the "dumbest" yet the "most authentic" method.
At the Agile Robots booth, an operator gripping the Diana3 master arm controls the remote Diana7 G2 slave arm, which mirrors every movement in real time. The system features millisecond-level bidirectional force feedback, faithfully transmitting sensations of contact, friction, and collision from a distance, while simultaneously capturing full-process multimodal datasets covering force, motion, and vision.
Star Era, meanwhile, is betting on the "whole body" approach. Its flagship humanoid, the Star Era L7, demonstrated coordinated teleoperation across 55 degrees of freedom, closing the technical loop between "whole-body data and model iteration."
Wearables — turning humans into "data collectors" — this year’s most crowded track.
Inex presented its EC-Gloves, capable of precisely reading joint angles across 20 degrees of freedom in the entire hand, supported by 1kHz high-frequency communication.

EC-Gloves data collection gloves; Source: Inex
Jianzhi Robotics has fully equipped the "head, hands, and body." Its Ego headband captures first-person perspectives, Dex dexterous gloves record fine hand movements, and full-body Mesh motion capture tracks movement trajectories—constructing a complete human data acquisition system.

gForce Ultra EMG wristband; Source: OYMotion
OYMotion’s gForce Ultra EMG wristband takes a different approach, using 8-channel dry electrodes to capture electrical signals from forearm muscles and arm posture.
Chaowei Power’s KAI Halo Lite headband combines four-camera vision with an IMU, enabling centimeter-level spatial positioning and automatic generation of action semantic labels. Orbbec, partnering with Ant Lingbo, showcased its EGO RGB-D series headset.
Tactile Sensing — from "ignored" to "contested ground."
Weitai Robotics unveiled three visuo-tactile data collection devices in a single sweep: the lightweight head-mounted VT-Ego, the handheld VT-TacFinger, and the VT-TacUMI-90 with a 90mm opening. Together, they form a full-chain tactile infrastructure spanning perception, execution, and learning.

XTac UMI G1 wearable gripper; Source: Qianjue Robot
Qianjue Robot’s XTac UMI G1 wearable gripper simultaneously captures five core dimensions of data: vision, touch, pose trajectory, and gripper status.
Yimu Technology has built a three-layer architecture centered on tactile sensing, spanning data collection, encoding, and multimodal large model training.
Body-Free and Simulation — using "smart force" to solve "hard problems" — is another emerging core route this year, breaking the limitations of lab scenarios.
Qiongche Intelligence’s RoboPocket system allows ordinary users to collect operational data without touching the robot itself. The company has already completed over 100,000 hours of data collection in home scenarios across 47 cities.
Guanglun Intelligence showcased a "data-evaluation-deployment feedback" continuous learning system for physical AI. Its SimFoundry simulation environment allows robots to repeatedly "practice" in a virtual world.
Kuawei Intelligence publicly unveiled its DexVerse™ generative simulation engine. Machine Science, meanwhile, centers on its self-developed differentiable multimodal physics simulation engine, RoboMirage, to build a "simulation plus video" dual data flywheel—cutting the cost of acquiring a single data point to just 1/20th to 1/200th of traditional methods.
Additionally, JD.com outlined a large-scale data collection plan leveraging its retail logistics scenarios. Chenjing Technology’s Looper kit features a "one-large, two-small" multi-view collaborative setup, while Lingdi Technology’s SynReal World full-stack data engine focuses on soft-body simulation.
With data collection solutions from over 20 companies competing head-to-head, this became one of the most crowded—and spirited—tracks at WAIC 2026.
Behind the 'Plethora of Schemes,' Just How Big Is the Data Gap?
Behind the blooming of diverse solutions lies an unsettling reality: the gap is alarmingly large.
At the WAIC "Wisdom Embodied Forum," Maoqing Yao, a partner at Zhiyuan Robotics and chairman and CEO of Mifeng Technology, dropped a figure that stunned the industry: the volume of data for physical AI is just one-twenty-thousandth of that for large language models. Qiangnao Technology offered an even more chilling estimate at the expo: as of early 2026, the total stock of high-quality embodied data globally stands at just 500,000 hours, while training a general embodied model requires upwards of 10 million hours—a gap exceeding 99%.
Pre-training data for leading international large language models has hit 100 trillion tokens—roughly equivalent to 10 billion hours of speech. Yet the information density of physical world data is far lower than text; one hour of video might contain less usable "physical knowledge" than a 1,000-word article. A gap of "over 10,000 times" separates the scale of embodied data from the training corpora of language models.
To put it in concrete terms: achieving usable embodied intelligence requires at least 10 million hours of multimodal interaction data. If the goal is a "ChatGPT moment"—where robots work "out of the box" and achieve a 70% to 80% success rate on common tasks—that figure needs to climb into the hundreds of millions of hours.
Yet the challenge isn't just scale; it's the cost of acquiring such vast quantities of high-quality multimodal interaction data.
Currently, real-robot teleoperation offers the highest precision and adaptability, but expensive hardware and low efficiency make it difficult to scale. Wearable collection strikes a balance with moderate costs and flexibility, yet the data deviates from the robot's own body, requiring complex domain adaptation. Simulation data offers high volume and low cost, but a natural "reality gap" exists—models trained in simulation often see significant performance drops when deployed on real hardware.
Consequently, scraping the vast amount of video on the internet has become a standard industry move. However, experts from Chenjing Technology argue that these videos hold limited value for embodied intelligence: "video without reliable pose data is just garbage." More precisely, what’s needed is "embodied experience" that carries spatial structure and task logic—stable trajectories, time synchronization, spatial alignment, and quality assessment. All are indispensable.
Focusing on hand manipulation, Chen Yilun, founder of Tashizhihang, noted at a July 17 roundtable that "public videos primarily offer geometric and visual information, whereas robot manipulation is a contact system. Relying solely on video makes it difficult to capture key physical quantities like force, touch, flexibility, and friction."
Consider the simple act of a human picking up a water cup. It looks effortless, yet it involves data across dozens of dimensions: visual positioning, arm trajectory, finger force, tactile feedback, and joint angles. Missing any single dimension significantly degrades the "experience" the model learns—like asking a blind person to mimic you drinking water using only sound; the result is predictable.

Source: Daimon Robotics
Therefore, Daimon Robotics argues that tactile sensing directly provides physical feedback—contact force, deformation, texture, and material—effectively compensating for the blind spots and illusions of vision. To this end, Daimon Robotics partnered with dozens of institutions to release Daimon-Infinity, the world's largest full-modal embodied dataset for the physical world that includes tactile data. It features 110,000 sensing units and high-frequency visuo-tactile sensors running at 120Hz.
Yet, based on current industry information, collection equipment, annotation standards, and coordinate systems remain fragmented across companies. Data formats are siloed, preventing heterogeneous data from being used for joint training, and in some cases, leading to "negative transfer"—where piling on more data actually degrades model performance.
Furthermore, insufficient scenario coverage and a severe lack of long-tail data remain critical issues. Most existing data comes from standardized lab settings—simple environments, regular objects, and repetitive tasks. The real physical world, however, is full of complex long-tail scenarios; data involving irregular objects, flexible materials, dynamic interference, and extreme conditions is incredibly scarce.
More critically, the industry prioritizes successful operation data while ignoring failures—yet failure cases are precisely what robots need to improve fault tolerance and generalization. This structural deficiency in scenario data directly causes models to "ace the lab but fail in the field," leaving generalization capabilities out of reach.
Summarizing these challenges, Maoqing Yao, partner and president of the Embodied Business Division at Zhiyuan Robotics, argues that scaling physical intelligence requires breaking through three walls: the Data Wall (scarce real-world interaction data and high acquisition costs), the Representation Wall (the lack of unified physical representation across tasks, scenarios, and robot bodies), and the Closed-Loop Wall (the high cost and slow feedback of real-world trial and error).
It is fair to say that no solution in the industry has yet overcome the trio of constraints: precision, scale, and cost. This is why data accumulation is lagging far behind the speed of model iteration.
From Hardware Stacking to Systemic Chain Building: Four Paths to Accelerate the Data Breakthrough
Facing this data bottleneck, this year's WAIC clearly reveals the industry's strategy: competition has evolved from comparing individual hardware components to vying on data systems and closed-loop ecosystems. Four paths have emerged as the industry's standard solutions: virtual-real integration, unified standards, distributed collection, and full-chain closed loops.
Virtual-real integration is currently the optimal way to balance data quality and scale. The industry has largely settled on a layered logic: "simulation for scale, real robots for precision, and wearables for diversity."
Source: Qiangnao Technology
Qiangnao Technology has built a three-layer data system: simulation and Ego video data for large-scale pre-training, wearable human data to enrich scenario diversity, and real-robot teleoperation data for high-precision fine-tuning—perfectly balancing cost, scale, and accuracy.
Guanglun Intelligence’s Real2Sim2Real loop and Mouxianfei’s "real-robot plus simulation" data factory both use real-world data to calibrate simulation environments, then leverage simulation to generate data in bulk, and finally verify and iterate on real robots. This process continuously narrows the reality gap, solving the issue where simulation data fails in real-world deployment.
Lingdi Technology’s soft-body simulation and Kuawei Intelligence’s industrial simulation generation further fill the gaps in special scenario data, continuously enhancing the industrial value of simulated data.
Data standardization and ecosystem alignment are key to breaking down data silos.
The push for standardization accelerated significantly at this expo, moving from company-specific standards to international and industry-wide collaboration. Guanglun Intelligence, the only Chinese firm involved, participated in drafting two international standards for global embodied simulation and data, leading the development of underlying data rules.
Kuawei Intelligence uses physical spatiotemporal alignment technology to enable the interoperable reuse of three heterogeneous data types—human, simulation, and real robot—reshaping the industry R&D paradigm to prioritize "effective data over sheer data volume."
Daimon Robotics’ Daimon-Infinity dataset employs a universal standardized format that supports adaptation across different robot bodies, making individual data points reusable and transferable, and significantly boosting data utilization efficiency.
The advancement of standardization has fundamentally ended the inefficient era of companies fighting alone and duplicating data collection efforts.
Distributed, body-free social collection is unlocking new sources of real-world scenario data. Traditional closed collection models have limited capacity, but the proliferation of lightweight wearable devices makes large-scale social collection feasible.

RoboPocket lowers the barrier to entry by combining wearable end-effectors with a mobile collection app; Source: Qiongche Intelligence
Daimon’s outsourced collection network spans over a dozen industries and hundreds of real-world scenarios, achieving an annual production capacity in the millions of hours. Qiongche Intelligence’s body-free collection solution lowers the barrier to entry, allowing ordinary users to contribute to home scenario data generation.
JD.com and Bodeng Intelligence leverage their own commercial and industrial scenarios to achieve "collection through operation and iteration through deployment." This allows them to continuously accumulate high-value real-world data without disrupting operations, effectively solving the problem of accessing complex scenarios.
A full-chain data closed loop builds a sustainable data flywheel. Data collection is not a standalone step; only by connecting the entire chain—collection, annotation, training, deployment, and feedback—can data achieve self-reinforcing value.
JAKA Robotics leverages its installed base of 30,000 robots to continuously collect real interaction data on industrial production lines, creating a positive cycle where "more deployment leads to more data, which leads to stronger models."
Weitai Robotics has established a three-layer hardware system for "perception, execution, and learning," achieving a full closed loop for tactile data from collection to application and learning. ZeroDimension and Bodeng Intelligence have built integrated platforms for collection, training, and deployment, forming a complete loop from data generation to application, fully unlocking the long-term value of data.
From 'Plethora of Schemes' to 'Sufficient Data': Three Mountains Remain to Be Climbed
While the industry's technical roadmap is now largely complete and the path to a breakthrough is clear, reaching the point of "sufficient data, usable models, and industrial deployment" still requires scaling three major mountains—making a large-scale breakthrough difficult in the short term.
The first mountain is industry-wide unified standards and ecosystem co-construction. The embodied data sector remains fragmented; collection norms, modality definitions, annotation systems, and evaluation standards are not yet unified, preventing data from flowing and being reused across companies, devices, and scenarios. The capacity of any single company is ultimately limited. Only through industry-wide collaboration can the 10 million-hour data gap be filled quickly. Although leading companies have initiated international and industry standard-setting, coverage is limited and progress is slow. Ecological mechanisms such as data rights confirmation, benefit distribution, and open-source sharing remain in their infancy, representing a core institutional barrier to speeding up the industry.
The second mountain is breaking the critical point of data cost and scaling. The cost of high-quality multimodal physical interaction data remains too high to support accumulation at the hundreds of millions of hours level. While simulation, automated annotation, and distributed collection continue to drive down costs, acquiring high-precision tactile, force, and failure-case data in real-world scenarios remains difficult and expensive. Industry consensus suggests that 2026 to 2027 could see the breaking of the 10 million-hour data threshold, enabling basic generalization capabilities. However, reaching the hundreds of millions of hours of effective data required for general embodied intelligence will still require 3 to 5 years of continuous iteration. The cost inflection point has not yet truly arrived.
The third mountain is the leap from "data volume" to "data quality." The industry is still in the early stage of "stacking data," with vast amounts of information suffering from homogeneity, low value, and a lack of differentiation. The core of future industrial competition will no longer be collection speed or data volume, but rather data quality, scenario diversity, modality completeness, and long-tail data coverage. Companies must evolve from simple hardware collection to a full-spectrum capability system involving data governance, screening, valuation, and iteration. This means precisely mining high-value data and eliminating ineffective redundancy to truly achieve a data-driven leap in model capabilities. This is a capability gap that the vast majority of companies have yet to bridge.
Conclusion
The data gap is vast, but vast does not mean unbridgeable. Large language models took a decade to leap from "unusable" to "ChatGPT," and the data infrastructure for embodied intelligence is evolving at a visible pace.
The outcome of this race will be decided by who can build the most efficient and sustainable data flywheel. Real-world scenarios generate data, which trains models, which are then deployed back into the real world to generate new data. The faster and more stable this cycle spins, the closer we get to "enough."
WAIC 2026 didn't give us the answer itself, but rather showed everyone answering the same question in different ways. When will the answer be revealed? Perhaps next year, perhaps in another three to five years. But the direction is clear, and the steps have been taken.
Models define the starting point, but data defines the endgame. This question of "enough" will ultimately be answered as the data flywheel accelerates. And we are witnessing that answer being written.








