XPeng Releases TuringViT for Smart Driving and Robotics
Automaker XPeng has released TuringViT, a vision encoder for smart driving, cockpit systems, and its IRON humanoid robot program.

Automaker XPeng has released TuringViT, a vision encoder designed for applications in smart driving, intelligent cockpit systems, and its IRON humanoid robot program.
TuringViT is intended for use in vision-language and vision-language-action models. XPeng offers the model in two versions: TuringViT-18L and TuringViT-24L. The company stated that TuringViT-18L, operating at a 1536x1536 resolution, achieved significantly higher throughput compared to previous benchmark models.
XPeng reports that TuringViT-18L's throughput was 3.04 times that of the Seed1.5-ViT model and 2.16 times that of the SigLIP2-ViT-L model. The model was trained on 850 million image-text pairs.
Despite its training dataset size, the model achieved an average score of 83.6% across six zero-shot benchmarks. This performance surpasses that of open-source baseline models, even those trained on up to 10 billion samples.