The Race for the Embodied "Brain" Is Not an Elimination Match

Edited by Yara From Gasgoo

The race for embodied intelligence is undergoing a clear shift in focus. While early industry battles centered on hardware and motion performance, the sector is now accelerating into a phase defined by the technology and commercialization of the embodied "brain."

This shift is accompanied by a surge in funding for "brain" technology companies, climbing valuations, and a continuous influx of new startups. The technical debate has moved past "does it exist?" to "whose solution is superior?"

World models have risen rapidly to become the industry focal point, thanks to their ability to anticipate environmental changes and their perceived potential for causal reasoning. This rise has cast doubt on the positioning of Vision-Language-Action (VLA) models, which previously powered most real-world applications, even sparking the radical claim that "VLA is dead."

But does the hype around world models truly signal that VLA is obsolete? And are world models, hailed as the key to generalization, really the ultimate answer for embodied intelligence?

These questions were at the heart of a recent discussion among industry experts at the "2026 Embodied Perception Fusion and Multimodal Large Model Innovation Seminar," hosted by Gasgoo Group. While their views weren't identical, they collectively mapped out the true current coordinates of the embodied "brain."

image.png

Image source: Gasgoo Embodied Intelligence

From VLA to World Models

The technological roadmap for the embodied "brain" is evolving. After two years of VLA dominance, the industry is now moving toward a multi-path parallel approach.

VLA, or Vision-Language-Action models, take multimodal observations like camera images and natural language instructions as input to directly output executable robot actions. They learn the entire process—"seeing the scene, understanding the command, taking action"—end-to-end within a single neural network.

At its core, this approach is a data-driven, end-to-end strategy rooted in imitation learning: it learns the mapping between "observation" and "action" from vast amounts of demonstration data.

In other words, VLA functions like a reflexive operator that acts on what it sees and hears. It learns to mimic "what an expert would do in this situation."

Because it avoids stitching together multiple modules, VLA offers clear advantages: shorter processing chains and faster response times, making it friendlier for edge deployment. UBTECH's Thinker-VLA, for example, is designed to run on the edge to reduce inference costs.

"But VLA has its limitations: it primarily generates action sequences and lacks an awareness of physical constraints and the ability to predict the next moment," noted Ji Haifeng, Director of Solutions at Ruieman Intelligent Technology (Beijing) Co., Ltd.

This means while VLA can answer "how to execute a movement stably," it struggles to answer "how the environment evolves after the movement." It particularly lacks long-term planning and causal reasoning capabilities, resulting in weak generalization beyond its training distribution.

This happens to be the inherent strength of world models: given a current state and a candidate action, they predict how the world will change next, answering "what happens if I do this," thereby enabling interaction with the physical world based on that reasoning.

So how exactly are "predicting the world" and "deciding on an action" realized? Jijia Vision offers this answer: the World Generation Model and the World Action Model.

"The World Generation Model predicts how the world will change after a given action. The World Action Model does the reverse: given current observations, it predicts the appropriate next action. Together, the generation and action models form a pair, creating a complete closed loop," explained Mao Jiming, Partner and Vice President at Jijia Technology.

Jijia Vision is a committed practitioner of the world model approach. Based on this logic, the company has established two major technical systems: the "World Generation Model" and the "World Action Model." The former includes GigaWorld for embodied intelligence, DriveDreamer for autonomous driving, and YiSu for content creation; the latter includes the general embodied brain GigaBrain and the world action model GigaWorld-Policy.

However, despite achieving a series of results in the world model field, Mao believes this technical route is still in its early stages of development.

image.png

Image source: Jijia Vision

At the 2026 World Robot Conference, Jijia Vision introduced its L1-L5 classification system for general world models:

L1: Generate visually realistic general video worlds from various inputs—making the world "look real" first;

L2: Generate efficient and reasonable general execution actions, bridging the gap from "seeing the future" to "making a choice";

L3: Generate geometrically accurate general spatial worlds, extracting geometric information from the pixel surface;

L4: Generate physically accurate general physical worlds, moving from spatial correctness to causal correctness;

L5: Generate general infinite worlds capable of long-term operation, allowing a world to exist long-term with continuous feedback.

"L1 and L2 correspond to the 'GPT-3 moment.' Once we enter L3 and L4, we believe world models will evolve into a second stage. At this point, the models can support agents in completing high-difficulty, long-horizon tasks requiring reasoning—akin to the current 'Claude moment.' This delivers powerful productivity and marks the entry into full-scale commercialization," Mao noted.

By L5, Mao expects the model to reach its culmination. At that stage, physical agents could achieve self-design and self-manufacturing, autonomously exploring the physical world to discover the unknown and ultimately achieving a state of self-evolution.

"We are currently focusing on L1 and L2 to achieve our near-term corporate goals," Mao stated.

New Paradigm Emerges: Reconstruction, Not Disruption

While the industry remains locked in debate over whether VLA or world models are closer to the endgame, another group of players is attempting to skip that choice entirely.

"I don't think the paradigm itself should be the core of the argument," said Lu Yao, Chief Scientist at Guangxiang Lab. "Instead, we should focus on the actual development needs of embodied intelligence and the differences between it and information intelligence. We need to find a development approach that truly fits—how do we better perceive the world, understand tasks, and ultimately give individual actions true universality and generalization? That is the ultimate goal of model development."

Whether the resulting model looks more like VLA or a world model is simply an outcome of development, Lu added. "We shouldn't limit model development from the outset by adhering to a specific paradigm."

This philosophy is central to Guangxiang Technology's development of "physical native intelligence." It is an approach of "starting with the end in mind"—reconstructing the model architecture of the embodied "brain" from the ground up.

image.png

Image source: Guangxiang Technology

According to Lu, Guangxiang's physical native intelligence emphasizes the needs of agent-environment interaction throughout the development process. It allows intelligence to emerge during interaction, enabling the system to understand how the world changes and the consequences of actions, all while adhering to underlying physical laws and safety constraints during learning.

"Humans derive only about 10% to 20% of their understanding and behavioral capability in the physical world from visual observation and imitation; the vast majority comes from perceptual interaction and trial-and-error feedback in the real world," Zhang Tao, Founder and CEO of Guangxiang Technology, previously noted.

Specifically, Guangxiang's physical native intelligence possesses three fundamental characteristics: decoupled state representation, temporal causal driving, and physical law constraints. These enable the model to perceive the world more clearly, understand problems more precisely, and ensure a more stable learning process.

It is worth noting that physical native intelligence still relies on the world model as its core and data-driven end-to-end training as its basic architecture. However, it delves deeper into the objective requirements of physical entities across three dimensions: the world model, data structure, and training algorithms. This includes addressing the scarcity of physical world data, the higher requirements for functional safety, and the pursuit of stability and longevity in model training.

Recently, Guangxiang Technology, in collaboration with Professor Li Shengbo's team at Tsinghua University, officially released its first-generation physical native world model, Phi-WM 1.0 ActEffect. According to reports, the model achieved leading scores on three international robotic manipulation benchmarks.

Beyond the physical native intelligence route, brain-inspired intelligence—neuromorphic intelligence—is emerging as another technological paradigm gaining industry attention.

Brain-inspired intelligence involves borrowing the neural mechanisms and cognitive architecture of the human brain to build intelligent systems. It allows robots to perceive, decide, reflect, and learn in a layered manner similar to a biological brain. Key players include Junaopanshi and Zhipingfang.

image.png

Image source: Junaopanshi

Junaopanshi recently released its first-generation brain-inspired cognitive world model, Cog-WM 1.0. While its core remains a world model, it distinguishes itself by adopting a brain-inspired JEPA (Joint Embedding Predictive Architecture) approach. Instead of generating future images at the pixel level, it predicts environmental states in an abstract latent space, using spatiotemporal memory to drive a unified architecture.

According to Junaopanshi, this allows robots to perform autonomous planning, semantic navigation, and manipulation in unfamiliar environments without pre-built maps or massive amounts of data.

Currently, Cog-WM 1.0 has demonstrated core capabilities on wheeled humanoid and quadruped robots, including autonomous navigation and path planning without pre-maps, spatiotemporal memory retrieval, and spatial relationship question-answering for object finding. It has also verified manipulation capabilities on wheeled humanoid robots, though it has not yet been deployed at scale commercially.

Zhipingfang's NeuroVLA, meanwhile, introduces the human brain's "cortex-cerebellum-spinal cord" synergy mechanism into the VLA architecture. The cortex handles semantic understanding and task planning; the cerebellum manages high-frequency motion coordination and dynamic correction; and the spinal cord is responsible for millisecond-level action execution and safety reflexes.

In late August, Zhipingfang's AlphaBot 2 (Aibao), powered by NeuroVLA, began serving cocktails and interacting with customers at a bar in Hong Kong's Lan Kwai Fong district.

In other words, even players championing "new paradigms" like physical native or brain-inspired intelligence still rely on world models or VLA as their foundation, rather than replacing them.

However, Wang Xianbin, a Partner and Vice President of the Gasgoo Research Institute, believes the rapid evolution of general large models from OpenAI, Anthropic, and domestic players is putting significant pressure on companies specializing in embodied "brains."

"The valuation bubble in this sector is relatively high right now," Wang assessed.

Yet, he also sees this as beneficial for companies focused on perfecting the "cerebellum"—the execution layer—particularly hardware manufacturers and those developing perception algorithms for vision and touch. "General model companies are unlikely to go very deep into the execution side," he argued.

Consensus is Fusion; Deployment is the "Hard Metric"

Although the technological roadmap for the embodied "brain" is far from settled, most leading players have already made their choice: don't rush to pick a side—focus on fusion first.

The reasoning is straightforward.

Asking a single model to both "think clearly" and "react quickly" defies common sense. You can take your time understanding a sentence or planning a task, but catching a suddenly falling part or recovering from a slip leaves only milliseconds for "thinking."

These differing demands for latency and computing power dictate that a single model architecture cannot dominate every scenario.

Ji Haifeng pointed out that for robots to actually be deployed in homes and other real-world environments, relying on a single model won't work. Instead, a "multi-brain fusion" is required to build the robot's "super brain."

"Perhaps newer technological paradigms will emerge in the future, but my current view is that these existing directions must converge," Ji stated.

While fusion is the consensus at the industry level, there is no standard answer for exactly "how" to fuse the models internally. Views differ on who handles the thinking, who handles the doing, and who provides the safety net.

One approach is to equip the robot with "fast and slow brains," each handling its own duties.

During the forum's roundtable discussion, Liao Yongxing, AI Algorithm Director at Fulaixin Material, was specific: the main brain handles high-dimensional work like semantic understanding, spatial cognition, and task planning; the cerebellum manages rapid responses like motion trajectory decomposition and physical interaction. This creates a layered structure of "slow understanding" and "fast reaction."

Another approach treats the world model as a "coach." On the training side, it generates data, offers predictions, acts as a referee, and performs physical verification, distilling these capabilities into the action model. The VLA acts as the "athlete," deployed for actual operation.

image.png

Image source: Jijia Vision

Take Jijia Vision's recently released GigaBrain-0.7. According to Mao, it deeply integrates a three-layer algorithmic pyramid: System 1, based on VLA, solves action and control problems; System 2, based on VLM, handles task understanding and planning; System 3 leverages world model capabilities for prediction and evaluation, providing the robot with the ability to "rehearse the future," thereby boosting the success rate and speed of difficult, long-tail tasks in real-world execution.

A third school of thought advocates training a single unified model that can both "anticipate and act," embedding world modeling, language reasoning, and action synthesis into one network.

Unitree Technology's UnifoLM-WLA-1.0, open-sourced in September, embodies this approach. Its highlight is the fusion of embodied reasoning, future dynamic region prediction, and discrete action learning within a single multimodal model. This synergistically enhances spatial perception, interaction prediction, and action generation, providing a unified multimodal representation foundation for future WLA training.

image.png

Image source: Galaxy General

Galaxy General's "Galaxy Star Brain" follows similar logic. Unlike the traditional three-tier separation of "perception-planning-control," where each layer is handled by different teams using different technical routes, the "Galaxy Star Brain" aims to connect these three layers into a single neural network, deeply fusing multimodal perception, long-term task planning, and real-time action execution within a unified system.

After all, stitching together multiple modules inevitably leads to information loss. A single model managing the process from start to finish theoretically offers a higher ceiling.

However, this path is the hardest to train and the least verified, remaining in its early stages.

Moreover, building an embodied "brain" with a higher ceiling faces challenges not just at the model level, but also from hardware constraints.

Xie Yipeng, Head of Unitree's Industrial Division, put it bluntly: high-compute chips are great, but their power consumption far exceeds what a robot's existing chips and modest battery capacity can handle.

In the smart vehicle sector, many models over the past two years have started carrying up to four NVIDIA Orin chips or even Thor chips, achieving compute power up to 1,000 TOPS. But fitting such high-power chips onto a humanoid robot presents a dilemma: a 1,000 TOPS Thor chip consumes over 120W, whereas the compute chips currently used in humanoid robots mostly consume between 5W and 20W. Given the very limited battery capacity of humanoid robots, it is difficult to support a compute chip drawing over 100W.

This raises another question: Should the "brain" of an embodied robot reside in the cloud or on the edge?

Currently, there is no standard answer to this question either.

However, the outlines of a division of labor are already clear: a hybrid of edge and cloud. High-frequency motion control stays on the edge, while large-scale training and long-term reasoning remain in the cloud.

If the "power wall" isn't broken, a hybrid deployment of "heavy models in the cloud, light strategies on the edge" will likely be the long-term norm.

Looking further ahead, Wang Xianbin believes that in the realm of true general large models, only top-tier tech companies, some automakers, and hardware manufacturers with sufficient resources and talent will likely succeed. Startups focused solely on building general large models will face significant pressure.

"However, for companies digging deep into very vertical niches, we believe there are still plenty of opportunities to win," Wang concluded.

Gasgoo not only offers timely news and profound insight about China auto industry, but also help with business connection and expansion for suppliers and purchasers via multiple channels and methods. Buyer service: buyer-support@gasgoo.com Seller Service: seller-support@gasgoo.com

All Rights Reserved. Do not reproduce, copy and use the editorial content without permission. Contact us: autonews@gasgoo.com

Related Documents(2)

Rankings of smart cockpit component suppliers in China (Jan.-Jun. 2026).zip
Rankings of ADAS component suppliers in China (Jan. - Jun. 2026).zip