Unified Agent: The Next Stop for Automotive AI

Edited by Taylor From Gasgoo

Gasgoo Munich-In 2026, the auto industry's AI narrative is getting a rewrite.

If 2024 was the year of "large models in cars" and 2025 belonged to "end-to-end," 2026 is staking its claim on a single phrase: the unified agent. It's the buzzword at nearly every tech launch.

Yet in production vehicles, the old split persists: cabin AI handles interaction, while ADAS handles driving. Is the industry's push for a "unified agent" a technical reality closing in—or just a narrative that needs a reality check?

At the fourth AI-Defined Vehicle Forum, hosted by Gasgoo, a roundtable tackled the question: "Full-Domain AI: Are Cars Moving Toward a Unified Agent?" Panelists Gao Jie, NIO's head of cabin AI; Wang Haowei, global ADAS chief at JOYNEXT; and Ma Jian, vice president of sales at THESEUS, hashed out the details.

When the industry says "unified," what exactly is it unifying? How do the technologies merge? How is safety maintained, and what serves as the foundation? These questions surfaced again and again.

The Unified Agent: What, Exactly, Is Being Unified?

Before answering whether cars are evolving into unified agents, a prerequisite must be cleared up: what does "unified" actually mean in this context?

A look at the actual solutions hitting the road reveals that "unified" is far more complex than it sounds.

image.png

Image credit: Geely Auto Group

Geely’s Full-Domain AI 2.0 hinges on a "1+2+N" multi-agent framework. The vehicle-level agent Eva sits at the core, coordinating two primary domain agents—driving and cabin—alongside sub-domain agents for chassis, energy, and body control. The goal isn't one model to rule them all, but rather a system where agents across domains can talk, negotiate, and collaborate.

Volcano Engine takes a "single brain" approach. Its end-to-end AI cabin architecture uses a central AI brain to deeply link the entire vehicle, closing the loop from perception to reasoning, execution, memory, and learning. Yet Volcano Engine vice president Yang Liwei draws a pragmatic line: "Volcano Engine is more like the brain, while smart driving is the cerebellum handling motion control." The brain interprets user intent and hands it off to the cerebellum—it doesn't replace it.

image.png

Image credit: NIO

NIO, meanwhile, redefines the boundaries of "unification" from the ground up. Its in-house full-domain operating system, SkyOS·Tianshu, bridges six domains—smart assisted driving, cabin, chassis, body, power, and cloud—unlocking over 1,600 atomic capabilities. It unifies management and scheduling across all domains. The significance? Rather than slapping a unified interface onto the application layer, it4 pulls capabilities from separate domains into a single scheduling1 scheduling system at the OS level—giving the unified agent a foundation it can actually run on.

image.png

Gao Jie, Head of Cabin AI at NIO

During the roundtable, Gao weighed in on what a vehicle-wide agent should look like. As touchpoints get smarter, agent-based software will multiply—"an inevitable form, an objective fact," he argued. Still, the user-facing interaction layer needs a unified agent in charge. "Since NOMI's first day, we wanted it to be the car's soul—like a human brain, aware of everything," he said. His ideal form for a vehicle-wide agent: "centralized yet democratic." Sub-agents run autonomously, but the master brain knows all.

Despite differing approaches, the industry's trajectory is clear: unify the interaction portal, unify the semantics, and execute safely in a distributed manner. "Unified" doesn't mean one model replacing all subsystems. It means preserving domain autonomy while letting interaction, semantics, and data flow freely.

For users, whether one agent or many hums in the background hardly matters. Seamless experience and reliable safety are the real benchmarks. And seamlessness hinges on whether the cabin and driving systems can truly work in sync.

Cabin-Driving Integration: How Far Has It Come?

With the "unification" direction coming into focus, the technical path gets more concrete: what logic should drive cabin-driving integration, and how far along is it today?

The industry has floated several solutions to the merging puzzle.

One classic route centers on chips and hardware architecture, with Horizon Robotics leading the charge. Its Xingkong 6P chip, built on a 5nm automotive process, delivers 650 TOPS of BPU computing power and 273 GB/s of memory bandwidth. By unifying memory and base software, it shifts vehicle computing from isolated domain controllers to central computing—halving space usage, cutting per-car costs by 1,500 to 4,000 yuan, and shrinking R&D delivery from 18 months to 8. "Integration is the immutable law of computing systems," Horizon founder and CEO Yu Kai argued.

image.png

Image credit: Dongfeng Motor

On the automaker side, Dongfeng Motor and Black Sesame Technologies co-developed the Tianyuan Smart Cabin Plus platform. Powered by the Wudang C1296 chip, a single processor simultaneously supports the smart cabin, L2+ driving assist, and parking—debuting in the Dongfeng eπ007, with broader rollout across multiple production models planned for 2026 and 2027.

Volcano Engine opts to bridge the gap at the model layer, without forcing chip-level unification. Yang offered a key data point: driving models demand inference frequencies of 10 to 48 Hz, while cabin models need just 1 or 2 Hz per second for smooth interaction. "Different processing frequencies multiply the demand for underlying compute. There's no need to merge the two models right now—the cost-performance ratio is poor." Volcano Engine's strategy: "precisely distill and summarize the user's intent for the driver agent, then transmit it to the driving system, closing the loop."

JOYNEXT takes a middle path, integrating chips and software via a central computing platform. Its nCCU series, built on Qualcomm's latest premium platform, supports deep fusion of the cabin and driving domains.

image.png

Wang Haowei, Global ADAS Head at JOYNEXT

For Wang, the real hurdle isn't technology—it's divergent needs. "Carmakers prioritize different things for the cabin and driving, and their software iteration speeds differ." His solution: "seek common ground while reserving differences." Unifiable requirements go into a unified architecture; the rest run on separate stacks.

NIO attacks the problem from the vehicle OS level. Its in-house SkyOS·Tianshu—the industry's first full-domain OS built for AI—operates on a core logic: cabin-driving fusion shouldn't be "connected" at the application or model layer. It must be hardwired at the OS level.

SkyOS·Tianshu bridges six domains—assisted driving, cabin, chassis, body, power, and cloud—unlocking over 1,600 atomic capabilities for unified management and scheduling across all applications.

This OS-level scheduling mindset shapes Gao's view of the boundary between driving and cabin. He offered a stark warning on the trend of driving domains encroaching on cabin interaction: "Over the past few years, smart driving has tried to do VLA and handle human-machine interaction from the driving domain. I think this direction is completely wrong. I totally disagree." His logic: driving decisions require a fast loop with "zero tolerance for misjudged frames," while cabin interaction is slow and high-level. The two are fundamentally different paradigms.

Approaches vary, but all paths must confront an unavoidable question: as cabin and driving fuse further, where do we draw the safety boundaries?

Interaction Can Unify; Safety Must Isolate

The deeper the technical debate goes, the more urgent safety becomes.

As large models stretch from the cabin into driving decisions, risk bleeds from the experience layer into road safety. Will large models make cars smarter—or just harder to control?

A March 2026 crash drew stark attention to this risk. A Lynk & Co Z20 was cruising when the voice assistant misinterpreted a command and switched off the headlights; the car then slammed into a highway median. The failure wasn't a lack of model "intelligence." It was the lack of isolation—interaction-layer output reached a safety-critical actuator unchecked.

images.jpg

Image credit: Waymo

Waymo offers a clear reference point. Its Ojai Robotaxi integrates Google Gemini as an in-car assistant, but under strict constraints: Gemini cannot alter the route, nor can it control windows or seats.

Crucially, Gemini is explicitly barred from claiming driving capabilities. One AI drives; another serves the passenger—and their permissions stay isolated.

On the hardware side, Horizon Robotics pioneered a "castle" architecture for physical safety isolation in its Xingkong chips. The cabin and driving domains run independently; the driving domain hits ASIL-D, the highest safety level, and a cabin reboot won't disrupt driving functions.

From a practical standpoint, Gao cut straight to the point: "In the short term, models are models, and product code is product code. All hard safety rules must be hardcoded. Don't hand them to the model." Wang was equally blunt: "No matter how smart autonomous driving gets, safely getting people from A to B is always the first principle."

The industry's de facto approach: unify at the interaction layer, isolate at the safety layer.

That doesn't conflict with the unified-agent vision—it's a realistic choice under today's engineering constraints. Yet it raises a deeper question: if safety isolation is mandatory and model layering is reality, what engineering foundation can keep a unified agent running sustainably?

Chips Matter, But What’s Truly Scarce?

Answering that requires returning to the basics of full-domain AI. Chip compute is growing, and central computing architectures are moving from concept to production. But whether these form the true foundation of full-domain AI remains a matter of debate.

Gao offered a clear verdict: "The core foundation of full-domain AI is full perception and full execution, plus central/pre-control architecture, the SDV software stack, and a data closed loop. If agents don't learn, the whole thing collapses." Chips and unified compute platforms, he stressed, "are bonus items, not the basics."

This view is gaining ground in 2026. From EV startups to legacy automakers, several have shifted their investment focus from features to underlying engineering systems.

Wang viewed it through the lens of AI infrastructure. In a keynote at the same forum, he argued that "models define the upper limit, but infrastructure determines AI's future." Future competition, he suggested, won't hinge on model capability alone—it will hinge on who can build the infrastructure to support continuous AI training, deployment, iteration, and scaling.

Though aimed at the broader AI industry, the takeaway applies to autos: as the unified agent moves from concept to production, what keeps it running may not be the model itself, but the engineering system that trains, deploys, and iterates it.

Image credit: @Grace Tao Lin-Tesla

Tesla's Shanghai Lingang AI Training Center, now online, closes the loop from data collection and local storage to domestic training, in-car deployment, and continuous iteration. This lets FSD handle full-pipeline data processing and model iteration within China, accumulating over 3 billion km of local road data. NIO, since 2021, has built its own data-loop framework and swarm intelligence validation system, enabling automatic data filtering on the car, cloud training, and parameter feedback. Its training architecture has evolved into a three-layer framework: world models, supervised fine-tuning, and closed-loop reinforcement learning.

These investments all point one way: making data drive model iteration more efficiently. That dictates not just today's feature performance, but whether the foundation can sustain continuous evolution.

Overall, the foundation of full-domain AI is shifting from standalone hardware or models to the engineering systems that keep turning data into model iterations. By 2026, this consensus has crystallized: chips and central computing remain necessary infrastructure, but the competitive focus is moving from "tops benchmarking" to "who can weave compute, data, and toolchains into an engineering system that keeps AI evolving."

The true scarce commodity is no longer any single technology—it's the engineering capability to keep technology running.

Conclusion

Returning to the roundtable's opening question: are cars moving toward a unified agent?

Technically, the answer is yes. Hardware solutions for cabin-driving fusion are in production, model-level semantic bridging is underway, and data-loop engineering frameworks are taking shape. But getting from "done" to "done well" spans a long road of engineering maturity, cost constraints, and safety responsibilities.

As for timelines, the panelists varied in specifics but agreed on direction. Wang figures comfortable interaction is "two or three years away," but deep integration into every vehicle layer—a true unified agent—needs "5 to 7 years." Gao bets "a point where the industry feels 'AI-defined vehicles' is valid could emerge in about two years," but "five years or longer before it becomes inevitable."

Beyond the timelines, the real question isn't "how long," but whether the unified agent can sustainably balance safety, experience, and iteration. Technical paths can advance in phases; safety boundaries must hold firm.

Gasgoo not only offers timely news and profound insight about China auto industry, but also help with business connection and expansion for suppliers and purchasers via multiple channels and methods. Buyer service: buyer-support@gasgoo.com Seller Service: seller-support@gasgoo.com

All Rights Reserved. Do not reproduce, copy and use the editorial content without permission. Contact us: autonews@gasgoo.com

Related Documents(3)

Rankings of smart cockpit component suppliers in China (Jan.-Jun. 2026).zip
Rankings of ADAS component suppliers in China (Jan. - Jun. 2026).zip
China Passenger Vehicle Export Rankings for H1 2026.zip