Gasgoo Munich- If the core battle in embodied intelligence last year centered on model architecture and demo effects, 2026 has shifted the industry's central conflict squarely onto "data."
Across the current market, leading players are scrambling to build data closed-loop systems. Publicizing dataset duration and expanding real-world robot collection have become industry norms—as if crossing the 100-million-hour threshold will naturally spark general embodied intelligence.
Yet as industrial implementation deepens, a growing number of frontline practitioners are realizing that data shortages are merely the industry's most superficial pain point.
At multiple forums during WAIC 2026, many experts agreed: beneath the anxiety over scale lie deeper structural issues—imbalanced data quality, fragmented standards creating data silos, and inefficient closed-loop iteration. These are the core bottlenecks determining long-term competitiveness and the true challenges in the deep end.
A Misread "Choke Point": Scale Is Just Surface-Level Anxiety
To achieve a "ChatGPT moment" for robotics, exactly how much data is required?
Yao Maoqing, partner and senior vice president at Agibot, president of its embodied business unit, and chairman and CEO of Maniformer, puts the figure at the 100-million-hour level.
In his view, the physical world is far noisier and more redundant. To enable robots to master general physical laws, understand open-ended instructions, and plan tasks for out-of-the-box deployment, accumulating 100 million hours of data is necessary just to achieve a 70% to 80% base success rate for common tasks.
Yet the industry is far from that target. Yao notes that pre-training corpora for leading large language models have reached 100 trillion tokens—roughly equivalent to 10 billion hours of speech. Compared to that, the total scale of current embodied data lags by a factor of over 10,000.
In reality, that 10,000-fold gap is merely the most visible "surface anxiety." What is truly stalling the deployment of embodied robots is a series of deep-seated misalignments hidden behind data supply.
The first is a structural mismatch between supply and demand.
The industry's current data bottleneck is not, at its core, a lack of absolute volume, but a severe shortage of high-quality, effective data that meets real-world deployment requirements.
"Data collected in labs is mostly pre-set, ideal scenarios—parcels are standard shapes," says Liu Lige, head of warehousing embodied robots at JD Logistics. "But in real warehouse environments, robots encounter reflective packaging, flexible bags, tangled tape, and numerous non-standard scenarios. It is precisely these long-tail situations that define the boundaries of a robot's capabilities."

Image Credit: JD Logistics
He draws a comparison from frontline experience: a packer working in a warehouse for one hour is several times less efficient than one working for a year. That efficiency gain comes from the latter's step-by-step iteration and familiarity with the scene through real operations. The logic of capability growth for robots is exactly the same.
Dong Le, executive vice president of the Beijing Institute of General Artificial Intelligence, shares this view: data must not only pursue scale but also guarantee quality and diversity, fully reflecting the long-tail task characteristics of the real world.
"If the model only repeatedly learns scenarios it has already seen, then no amount of data volume is meaningful," Dong notes. After all, the ultimate goal of data collection is to enable the model's intelligent growth. If it stays stuck polishing familiar scenarios, the model's generalization capability is out of the question.
The second is a misalignment of standards across cross-entity collaboration.
The embodied intelligence industry has yet to form unified data format standards. Data definitions, labeling norms, and storage formats vary wildly across different robot manufacturers and model teams, making cross-entity data reuse extremely difficult. Widespread redundant construction and "reinventing the wheel" are common phenomena.
Liu Dong, CEO of XYZ Embodied AI Co., Ltd. , revealed that early on, the company attempted to train on a mix of over 200,000 data points from more than 20 brands. Without standardized processing, directly mixing the raw data for training resulted in a "mess" that was simply unusable.
"Later, to unify data from these 20-plus robot types, we built our own data collection and processing platform to standardize all collection perspectives and robot joint information into a single format. Only then was training feasible," Liu said. He suggests that when collecting data across different robots, a system must be designed to be compatible with various robot data, rather than just taking raw materials from manufacturers.
Moreover, even for the same robot model, different production batches can have significant tolerances given current manufacturing levels—posing even higher demands for mixing real-world robot data.
The third is a misalignment in the closed-loop of production and R&D iteration.
For the embodied intelligence industry, an efficient data closed-loop is the core path to achieving continuous model evolution, building competitive barriers, and ultimately moving toward large-scale commercialization.
But in reality, the "data closed-loop" touted by many companies remains at a conceptual level, with mid- and back-end data governance and feedback pipelines not yet truly operational.
"Look at autonomous driving data pipelines: they have very complex pre-labeling, quality control, and post-labeling stages," notes Wu Wei, CEO of Manifold AI. "But in embodied data collection, the focus is now heavily on front-end acquisition, while the middle and back-end stages of data cleaning, labeling, and quality control are missing."
Hu Weiqi, head of commercialization for MiniMax China, also believes the core problem for the industry now is improving efficiency—both in data generation and usage. Crucially, it involves how to loop failure cases back into the model and the robot after they occur in real-world scenarios to achieve capability improvement.
Thus, while the 10,000-fold scale gap is stark, three structural shortcomings—imbalanced data quality, fragmented standards creating silos, and inefficient closed-loop links—are the far more difficult barriers standing in the way of large-scale deployment.
Consensus on the Solution: Data Is Not Piled Up, It Is Blended
Since the core challenge for embodied data goes far beyond scale, the solution naturally cannot be limited to the crude model of "stacking hours."
The industry has reached a basic consensus: no single data collection method can cover all needs. Data from different sources, costs, and precision levels each has its role. The evolutionary path with the highest return on investment lies in scientifically blending them to form a complement.
More specifically, it is widely believed that the demand for different data types in embodied intelligence presents a typical "pyramid" structure: from bottom to top, it consists of public internet video data, non-robot collection data, and real-world robot testing data.
Internet video data at the base of the pyramid is characterized by massive volume, low acquisition cost, and broad scenario coverage. However, its quality is uneven, and the efficiency of transferring actions to robot bodies is limited. Thus, it is primarily used for the general pre-training stage to help the model build a basic understanding of the physical world.

Image Credit: China Telecom AI Research Institute
The middle layer—non-robot collection data, including data acquired via UMI or first-person perspectives—features high structuring and flexible scenarios. It is currently the optimal solution for balancing cost and effectiveness. However, its shortcomings lie in insufficient precision for complex, fine operations and a certain gap in robot adaptability.
The real-world robot teleoperation data at the pyramid's top comes directly from the target robot body. Among the data types, it boasts the highest precision and best scenario adaptability, and it is the key to pushing model capabilities from "usable" to "useful."
"Large models do not learn isolated task skills from data; they learn an underlying understanding of real-world physical laws, task logic, and causal relationships," Dong says. "In this regard, real-world robot data can provide genuine verification and calibration, thereby effectively improving the model's generalization capability."
Zhu Senhua, CEO of EBKernel, even goes so far as to call real-world robot teleoperation datasets the "gold standard" for embodied training.
Yet, even with its highest value, real-world robot data has unavoidable limitations: narrow scenario coverage, slow scaling speed, and persistently high collection costs. It is difficult to rely solely on this data to complete the pre-training of general capabilities.
Precisely because different data types have their own pros and cons, the industry universally adopts a multi-source training strategy.
Liu revealed that when XYZ Embodied AI Co., Ltd. trains its embodied interaction world models, non-robot data accounts for about 90%, used for basic pre-training, while real-world robot data makes up about 10%, used for post-deployment training alignment.
"Although real-world robot data accounts for only 10%, it is indispensable. Once the pre-trained base model is complete, to actually deploy it on a robot, you must use data collected from the target robot for post-training; otherwise, the model cannot complete transfer and mapping."
It is worth noting that, beyond the data acquisition methods mentioned above, synthetic simulation has also attracted widespread attention due to its lower cost and flexible scenario expansion capabilities. However, because a gap between the virtual and real environments always exists in simulation data, its role varies significantly across different companies' technical roadmaps.
Wu Wei explicitly stated that his team has completely abandoned the pure simulation data route.
The underlying logic is this: a simulator is essentially a small model, and using data generated by a small model to train a large model naturally creates a capability ceiling. "The pure simulation route iterates very quickly in the early stages, but in the later stages, it will inevitably hit a distinct performance ceiling. It can even be difficult to pinpoint which part of the data is causing the problem. This is the core reason we abandoned pure simulation data."
However, many viewpoints hold that simulation can serve as a crucial tool for data augmentation and evaluation, playing an irreplaceable role.
Zhu Senhua stated that the core value of the simulation route lies precisely in data amplification. "We do not advocate over-reliance on simulation data, as the gap between the virtual and the real objectively exists, but we have high hopes for the simulation route. If we can achieve breakthroughs in simulation toolchains with high physical, visual, and interaction realism, it will greatly alleviate the data supply pressure for the entire industry."
Sun Jiaqi, co-founder of Motphys, also pointed out that simulation's greatest advantage is that no human intervention is required during the data augmentation and generation stages—computing power directly determines data output. "It can conduct controlled variable experiments on the causal chains of tasks, precisely pinpoint the model's capability shortcomings, and supplement data in a targeted manner—something that is very difficult to achieve with real-world robot collection."
Beyond these data categories, some "atypical high-value data" that the industry has long underestimated are also gradually becoming a focus of attention.
The first category is failure and anomaly data.
"We found in data collection that the value of failure data is severely underestimated by the industry," Liu said. "In the VLA era, everyone kept only successful cases and discarded failure data directly. But after upgrading to the world model paradigm, failure data is equally valuable. It allows the model to learn where the boundaries of a task lie and what actions lead to failure, thereby accumulating experience from failures to ensure the success rate of final execution."
Liu Lige also believes that robot failure data holds extremely high value. "Especially in real-world scenarios, such as when a robot drops a package or gets stuck, the recovery data—whether a human intervenes via teleoperation or corrects it in other ways—is incredibly valuable to the robot because it precisely fills the gaps in the current strategy."
For this very reason, many teams instruct data collectors to intentionally perform failed actions during collection.
The second category is scenario-native business attribute data.
"Many people overlook that real business scenarios inherently come with massive amounts of ready-made labeled data," Liu Lige offered. In warehouse scenarios, information such as package weight, dimensions, material hardness, and category attributes is fully recorded in the Warehouse Management System (WMS). Jointly training this business attribute data with visual and action data can significantly improve the model's precision in understanding object characteristics.
This means that competition in the embodied data field is quietly shifting from a battle over "collecting more" volume to a contest of systematic capabilities in "collecting precisely and blending well."
Four Types of Players Enter the Fray—Who Will Seize Data Dominance?
As the strategic status of data in the embodied intelligence industry continues to rise, players from different backgrounds are entering the fray, building their own capability systems around data collection, governance, and closed-loop construction. Overall, there are currently four core types of players in this sector.
The first category comprises full-stack robot manufacturers that hold the hardware and deployment scenarios, represented by companies like AgiBot and Galaxea AI. Their core advantage lies in the ability to define data standards from the hardware level. Under the real constraints of edge computing power and control frequency, they can explore optimal collection and iteration schemes to achieve collaborative software-hardware optimization.
The second category consists of pure embodied model companies taking a light-asset route, including RoboScience and Manifold AI. These companies focus primarily on the development of foundational embodied models and do not engage in hardware manufacturing. Their core advantage is a deeper understanding of model training needs, making them, to some extent, the natural definers of industry data standards.
"The algorithm paradigm directly determines data format and labeling requirements," Zhu Senhua argues. "Only teams that truly understand the underlying logic of algorithms and the pain points of model training can define scientific and efficient data collection norms and standards."
The third category includes industrial players that hold real business scenarios, represented by JD and China Mobile. These players naturally possess normalized, real-world business scenarios and can achieve "non-stop data collection"—frontline workers can complete data collection synchronously during normal operations, directly eliminating the labor cost, which is the highest expense in the data collection process.
Wu Wei predicts that the endgame of data operations will most likely land in the hands of these scenario owners who possess large-scale offline manpower, as their inherent cost advantage is difficult for pure technology companies to match.
The fourth category is third-party data service providers, represented by Maniformer. These companies position themselves as the "utilities" infrastructure for the entire industry. They neither develop models nor produce robots, focusing instead on providing full-link data services ranging from collection hardware and labeling governance tools to standardized datasets and evaluation systems.

Image Credit: Maniformer
Because these four types of players have distinct endowments and complementary advantages, different judgments are forming within the industry regarding who will ultimately hold data dominance.
Liu Dong believes the construction of future data training grounds and data factories will proceed in two steps: model companies will first build medium-scale factories to validate data collection paradigms, then complete large-scale data production through industry collaboration.
"Pure third-party service providers do not understand the underlying architecture and data paradigms of the models, so the data they produce is difficult to match directly with training needs. Conversely, model companies find it hard to synchronize every detail with third parties. Therefore, model companies must first build medium-scale data collection factories to polish the complete paradigms and production processes for data collection, cleaning, and labeling themselves, even using model methods to assist in data processing," Liu said.
Once standardized SOP production flows are formed, the standards can be abstracted and exported for collaboration with external partners. "For example, connecting with a partner like JD that has large-scale manpower and can conduct real-world robot collection, we define the basic data collection specifications, then bring the raw data back to our own data factory for finishing, ultimately using it for model training."
Zhu Senhua shares this view: the construction of data factories and the formulation of standards and SOPs must inevitably be led by model companies.
Wu Wei, starting from a cost perspective, believes that the endgame dominance of the embodied data industry will most likely be held by large-scale manpower operators, such as Meituan and JD.
"Their cost advantage is too obvious. Delivery riders and logistics workers are already working on the job every day, effectively serving as non-stop data collectors, which directly eliminates the highest labor cost in data collection." Accordingly, he is not optimistic about pure model companies building overly heavy data operation pipelines or recruiting large numbers of people for collection.
"However, model companies must firmly control the right to define data standards. Core rules like sensor selection and data precision requirements must be kept in their own hands," Wu said.
From this perspective, it is highly unlikely that the future embodied data industry will see a single winner-take-all scenario. Instead, it will move toward a layered, collaborative industrial ecosystem where each player deepens their expertise in their respective advantage tracks, jointly driving the maturity and perfection of the data system.
Conclusion
From "competing on hours" to "competing on quality," and from "fighting alone" to "division of labor," the data race in embodied intelligence is quietly undergoing an upgrade.
Two years ago, the industry was still debating whether embodied intelligence should pursue a big data route. A year ago, everyone was still competing over who had the larger volume of public datasets. Today, practitioners have already begun to delve into more systematic propositions such as standard unification, structural ratios, closed-loop efficiency, and industrial division of labor.
This shift in focus is, to some extent, a signal that the track is maturing. After all, the ultimate value of data has never been about simple numerical stacking, but about enabling robots to stably create value in the real world. When all exploration into data ultimately points to deployment and industrial value, that is when the race has truly entered the meaningful deep end.









