Gasgoo Munich- 1 million hours isn't enough.
On August 31, Maniformer—a subsidiary incubated by Zhiyuan Robotics—announced that the 20,000th unit of its MEgo series had rolled off the production line. At the same time, it released a dataset comprising 1 million hours of embodiment-agnostic data. This batch covers 22 major scenario categories, more than 500 task types, over 10,000 real-world scenarios, and upwards of 50,000 object categories.

Image source: Maniformer
Just six months after its founding, Maniformer has already become the first company in the industry to publicly claim a production capacity for embodiment-agnostic data at the million-hour level.
Yet on the same day, Yao Maoqing, Maniformer's chairman and CEO, set the industry's next target at a completely different order of magnitude. “To reach AGI, physical AI will ultimately need tens of millions, even hundreds of millions of hours of data. Today is just the starting point.”
The jump from 1 million to 100 million represents a 100-fold increase.
That vast gap evokes a familiar sense of déjà vu.
Over the past few years, the large language model industry has already lived through a similar "scaling" narrative: more data, larger parameters, more computing power—and then waiting for capabilities to suddenly emerge at a critical threshold.
Now, that logic appears to be transplanted to robotics.
Consequently, data collection factories are proliferating, "no-body" devices are entering mass production, and real-world data is being measured strictly in "hours"—with targets of 10 million or even 100 million hours following fast on the heels of the first million.
But a more fundamental question deserves asking first: Does embodied intelligence actually know what kind of data it needs for those 100 million hours?
If that question remains unanswered, does "1 million hours" really mean robots are a step closer to general intelligence? Or does it simply mean humanity has, for the first time, acquired the ability to mass-produce physical-world data?
These are not, in fact, the same thing.
Large Model Scaling Laws Don't Simply Apply to Robots
It is not hard to understand why robots lack data.
Language models enjoy a natural advantage: over the past few decades, humans have already digitized vast amounts of knowledge.
Web pages, books, forums, code, images, and videos are, in themselves, ready-made training materials.
But robots ultimately have to solve problems in the physical world.
How to pick up a cup, how to fold clothes, how to place items of different materials on a shelf—these operational skills that humans take for granted do not naturally exist on the internet.

Yao Maoqing notes that the scale of data for physical AI is still "four or five orders of magnitude" smaller than that for language models. In his view, a robot's true skills come from experience: "Only after seeing it can one do it, and only then can one generalize."
The problem lies here: There is a vast distance between "robots lack data" and "robots need 100 million hours of data."
A critical foundation for the scaling of large models is that data like text and images have relatively mature tokenization methods, while model architectures, training objectives, and evaluation systems have gradually converged.
The physical world, however, is far more complex.
Take the simple act of "picking up a cup." Different robots might use two-finger grippers or five-finger dexterous hands; different bodies have varying degrees of freedom, dimensions, torque, and control frequencies. Data generated by a human performing an action cannot necessarily be directly converted into control signals for a robot.
More importantly, today's embodied intelligence hasn't even fully answered a basic question: What data is actually most useful?
Is it better to fold clothes once in 1,000 different households, or for one robot to fold clothes 100,000 times?
Is a human's first-person video more important, or are the robot's joint trajectories more critical?
Is it more important to expose the model to more tasks, or to ensure it performs a smaller number of tasks with sufficient reliability?
There are currently no unified answers to these questions.
Therefore, at least for today, "100 million hours" cannot be understood as an engineering conclusion similar to "a model requires at least X tokens for training."
It represents more of an industry judgment: If embodied intelligence ultimately needs to achieve generalization in the open world, then today's data scale is far from sufficient.
As for exactly how much is lacking, the industry is still figuring that out.
Why Is embodiment-agnostic data Taking Off Right Now?
It is precisely against this backdrop that embodiment-agnostic data has begun to heat up rapidly.
In the past, the most direct way to acquire robot data was to let the robot do the work itself.
Operators would control robots via teleoperation to complete tasks like grasping, moving, and organizing. What the robot saw, how its joints moved, and the trajectory of its end-effector were all recorded simultaneously.
The advantage was that the data was naturally native to the robot.
The drawback was equally glaring—production efficiency was too low.
He He, vice president of JD Technology Group, recalled at the event that a year ago, during physical-world data collection, an operator might work with a robot for 8 hours and yield only 1 hour of effective data—all while absorbing the costs of the robot, electricity, labor, and venue.
Yao offered another figure: today, tens of thousands of hours of real-robot data is a common requirement for some models. Acquiring that much data could require thousands of robots collecting continuously for six months.
This exposes a glaring contradiction.
On one side, the industry is starting to believe that data should scale;
On the other, traditional data production methods simply cannot scale.
embodiment-agnostic data is essentially a solution to this contradiction.
Humans no longer need to teleoperate a physical robot. Instead, they can wear first-person or wrist devices, or simply hold a gripper to complete tasks directly.
Of the 1 million hours of data released by Maniformer this time, 500,000 hours came from bare-hand collection, 350,000 hours from wrist-mounted devices, and 150,000 hours from gripper tools.

The comparison in the on-site cafeteria was even more intuitive.
Traditional teleoperation might yield 1 hour of effective data in an 8-hour shift. In contrast, catering or cleaning staff wearing devices during their normal workday could generate 5 to 6 hours of data in the same period.
Therefore, the emergence of "no-body" technology does not primarily solve the problem of "how to make robots smarter."
It solves a more practical problem: how to lower the cost of producing physical-world data.
These are two completely different issues.
Being able to cheaply produce 1 million hours of data does not guarantee that this data will deliver improvements in model capabilities proportional to its volume.
This is the most critical distinction to make when viewing this "1 million hours" milestone.
It proves data production capacity first, not model capability.
The Greatest Value of 1 Million Hours May Not Be the "1 Million"
In fact, even Maniformer itself hasn't placed all its bets on the number of hours.
The 1 million hours of data released this time was accompanied by emphasis on its diversity: 22 major scenario categories, over 500 task types, more than 10,000 real-world scenarios, and upwards of 50,000 object categories.
This actually reveals a key difference between scaling in embodied AI and scaling in large models: Physical-world data is difficult to measure solely in "hours."
Two datasets could both be called 1 million hours.
One might consist of massive amounts of repetitive tasks, while the other covers completely different environments, objects, and operations. Their significance for a model is clearly different.
Similarly, within a 30-minute data clip, there may be only a few minutes of truly valuable operations.
Therefore, while the hour count is the easiest metric to understand, communicate, and track across the industry, it may not be the most accurate gauge of the value of embodied data.
This is why, when seeing figures like "1 million hours," "10 million hours," or "100 million hours" today, it is necessary to maintain a certain perspective.
They indicate that supply scale is expanding rapidly, but they cannot be directly equated to improvements in robot capability.
What really needs to be observed is a different curve:
When data increases from 100,000 to 1 million hours, how much does the model actually improve?
Has the success rate increased?
Has generalization capability improved in unseen environments?
If the robot body is swapped, how much of the previously learned capabilities are retained?
Have tasks that were previously impossible been learned thanks to this data?
Without corresponding answers to these questions, data scale itself risks becoming a flashy yet hollow industry KPI.
At least at this launch, Maniformer focused on showcasing data scale, coverage, and collection and governance capabilities. As for how much capability gain the 1 million hours of data might bring to a specific model, the event did not present a complete set of results directly correlating with that data volume.
Therefore, the "industry's first million hours" is worth recording, but it is not yet a technical proof on par with "capabilities emerging after large model parameters cross a certain threshold."
This distinction needs to be made.
Do Robots Need to "See More" or "Practice More"?
Taking a step further, embodied intelligence faces a problem even more fundamental than data volume.
Human learning is not simply a matter of replicating every action ever seen.
Once a person learns to open a specific type of drawer, they will likely still know how to "pull open a drawer" even if they change rooms, handles, or heights.
This is ultimately the capability robots are striving for.
Generalization.
In this sense, the value of large-scale real-world data is easy to understand:
A lab can build one restaurant, but it is hard to build 1,000 different ones.
In the real world, lighting, table height, object placement, foot traffic, and distractions change every moment. The more environments a robot sees, the greater its theoretical chance of learning the more universal laws underlying tasks.
This is also the most attractive aspect of embodiment-agnostic data.
It extends data sources from limited robot labs to the real world.
But this still only solves half the problem.
Seeing enough is not the same as doing enough well.
The human body is different from a robot.
Human operational experience can tell a model "this is how you hold a cup," but it cannot naturally guarantee that a specific robot knows how many degrees its joints should rotate or how much force its gripper should apply.
Therefore, embodiment-agnostic data and real-robot data are not in a simple substitution relationship.
Yao also admitted during a group interview that embodiment-agnostic data is primarily used for pre-training and general representation learning. When it comes to post-training and deployment for specific robots and specific tasks, "you definitely cannot get around real-robot data from the corresponding body."
Zhu Yajuan, head of product ecology at Tencent's Robotics X Lab, was equally cautious. She noted that the "heterogeneous + gripper" approach of embodiment-agnostic data can "to a certain extent" replace problems previously solved by teleoperating a physical body. However, the alignment of fine movements during cross-body transfer still requires real-robot data for calibration.
These two qualifiers are actually more noteworthy than "1 million hours."
Because they show that the industry still hasn't found a single type of data that can solve everything.
Robots need to both "see a lot" and "practice a lot."
And exactly how to combine these two needs is still in the exploratory stage.
The "Scaling Law" of Embodied AI Has Not Been Proven
So, returning to the initial question: Is 1 million hours not enough to "feed" the robot?
The answer might be that—at least today—there is no answer to that question. But what is truly worth noting is something else: the industry has begun to use "data scaling" to answer the question—and scaling itself may well become part of the answer.
Yao believes it will take hundreds of millions of hours to pre-train a sufficiently good embodied foundation model.
This judgment is not without logic.
If the goal is to deploy robots into homes, malls, factories, warehouses, and other open environments, the combinations of objects, tasks, and environments they will face are nearly infinite. Existing data clearly falls far short of covering them all.
Yet, between 1 million and 100 million, there is no verified Scaling Law.
Will a capability leap occur at 10 million hours?
What is the essential difference between 50 million and 100 million?
As model architectures improve, might data demand actually decrease?
In what proportions should different types of data be combined?
There are currently no definitive answers to these questions.
Therefore, the more accurate significance of this 1 million hours of data may not be to prove that "robots are 1% closer to AGI."
AGI is not a progress bar that can be simply calculated by data hours.
What it truly demonstrates is something else:
Physical-world data is, for the first time, becoming capable of industrial-scale production.
Twenty thousand collection devices, 1 million hours of data, real-world scenario collection, and increasingly professional data processing pipelines all point to the formation of a new industry—human physical operational experience, previously unrecorded, is beginning to be systematically digitized.
As for the extent to which this data will ultimately drive the scaling of robot capabilities, that remains a question for model training and real-world deployment to answer.
Therefore, for embodied intelligence today, the real danger is not too much data, but prematurely equating "expanding data scale" with "improving intelligence capabilities."
After 1 million hours, the industry will naturally continue to chase 10 million and 100 million hours.
But what needs to be proven in the next stage is not just whether the numbers can continue to grow.
It is whether robots are actually becoming smarter as a result.
In other words: The industrialization of data production capacity and the emergence of intelligence capabilities are two independent events. The former is happening; the latter remains to be proven. Conflating the two is precisely the part of the current embodied AI narrative that warrants the most caution.
Of course, before proving that, there is an even more practical question.
If the industry is truly preparing to march toward 100 million hours, where exactly will the remaining 99 million hours come from?
That is the question we will discuss in the next installment of this series.









