XPENG Launches TuringViT Visual Encoder, Reconstructs Training Paradigm for Visual Large Models in Physical AI Scenarios

Edited by Taylor From Gasgoo

Gasgoo Munich-XPENG has officially launched the TuringViT efficient visual encoder. The system systematically reconstructs the architecture design, data paradigm, and training workflow of visual encoders for the VLM/VLA era. It is set to underpin three key business scenarios: smart driving, smart cockpits, and the IRON humanoid robot.

10579446ca07c9b824ca9b71abe65fa0.png

Image source: XPENG

The industry's current reliance on open-source generic ViT solutions faces a triple bottleneck in real-world scenarios involving high resolution, multi-view angles, and continuous video frames: computational cost, data efficiency, and scenario adaptation. TuringViT addresses this by advancing simultaneously across architecture, data, and training, proposing a new technical path.

Architecturally, TuringViT uses Turing Linear Attention as its core computing unit, building a hybrid Turing Block structure. Linear attention handles the heavy lifting for global context aggregation, reducing computational complexity from quadratic to near-linear. Tests show that at 1536×1536 resolution, TuringViT-18L achieves inference throughput 3.04 times that of Seed1.5-ViT. The system is available in two versions: 18L and 24L.

On the data front, TuringViT introduces the VISTA-Curation multimodal data governance pipeline. By optimizing multi-stage filtering and annotation for image-text and video data, it enhances the supervision value of individual samples. The system completed training using only 0.85B image-text pairs—roughly 10% of SigLIP2-L's training data scale—achieving an average accuracy of 83.6% across six zero-shot classification benchmarks, including ImageNet-1K.

For training, TuringViT employs a four-stage progressive native dynamic resolution paradigm. It adapts to downstream VLM/VLA input characteristics starting from the pre-training phase and, paired with 2D rotary position encoding, supports inputs of varying sizes and aspect ratios.

In application, TuringViT will serve as the core visual encoder for XPENG's second-generation VLA model, processing multi-channel surround-view camera inputs. It also supports vision-language model alignment in smart cockpits and provides foundational perception capabilities—such as object recognition and spatial relationship understanding—for the IRON humanoid robot. XPENG noted that the system's architecture and training workflow do not rely on specific hardware platforms, offering the industry a reproducible path for training large visual models.

Gasgoo not only offers timely news and profound insight about China auto industry, but also help with business connection and expansion for suppliers and purchasers via multiple channels and methods. Buyer service: buyer-support@gasgoo.com Seller Service: seller-support@gasgoo.com

All Rights Reserved. Do not reproduce, copy and use the editorial content without permission. Contact us: autonews@gasgoo.com