XPENG has launched TuringViT, a vision encoder designed for vision-language and vision-language-action models, applicable in smart driving and its humanoid robot program, IRON. The company offers two versions: TuringViT-18L and TuringViT-24L, with the former achieving a resolution of 1536×1536 and impressive throughput metrics.
The significance of TuringViT lies in its performance, with TuringViT-18L reportedly reaching 3.04 times the throughput of Seed1.5-ViT and 2.16 times that of SigLIP2-ViT-L. This advancement is attributed to the model being trained on a substantial dataset of 850 million image-text pairs, resulting in an average score of 83.6% across six zero-shot benchmarks, surpassing open-source models trained on 10 billion samples.
Looking ahead, the implications of TuringViT for smart driving and humanoid robotics are noteworthy, as XPENG continues to innovate in these fields. No further timeline was disclosed at the time of publication.
Editor's Note
XPENG's introduction of TuringViT highlights the growing importance of advanced vision encoders in the robotics and automotive sectors. As companies increasingly integrate AI-driven technologies into smart driving and humanoid robotics, the competitive landscape is likely to evolve, emphasizing the need for high-performance models that can process vast amounts of data efficiently.
Leave a comment