Gasgoo Munich- XPENG officially launched the full rollout of its second-generation VLA system, XOS 6.3.0, on September 22. The core of this upgrade lies in the physical world foundation model, which now incorporates "time" into its framework—signaling a shift from static 3D spatial recognition to dynamic 4D space-time understanding.

Image source: XPENG
According to XPENG, the new version introduces the Infini-VLA long-sequence architecture, capable of retaining 30 seconds of road environment data. It also features the X-Foresight world model, which predicts the movements of surrounding traffic participants up to six seconds in advance. Edge model parameters have expanded by 3.5 times, boosting end-to-end response speeds by 300%.
A Lite version of the second-generation VLA is rolling out simultaneously, bringing upgrades specifically to the Turing Max model. Distilled from the same high-end architecture, XPENG utilized learned token compression and distillation training to deploy the foundation model's capabilities on platforms with lower computing power. Chairman He Xiaopeng revealed that the distilled Lite model has entered mass production, with real-world tests placing it among the top tier of domestic L2 autonomous driving systems.
Functionally, XOS 6.3.0 adds options for nose-in and tail-out parking, allowing drivers to freely switch parking directions after selecting a spot. It also supports voice-controlled parking exit. Inside the cabin, the update introduces split-screen interaction and CarPlay, enabling multiple applications to run simultaneously on the same display.
Industry observers note that automakers are exploring the reuse of a single large model foundation across vehicles and robots to amortize R&D costs. However, large-scale commercialization of such technologies still demands sustained real-world verification.









