Senior Vision-Language ML Engineer, Autonomous Driving

XPENG

Santa Clara (CA)

On-site

USD 175,000 - 296,000

Full time

47 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Massive real-world data
Industry-scale compute
Top-tier researchers and engineers
Competitive compensation
Snacks, meals, and team events

Job summary

XPENG is seeking a Machine Learning Engineer / Research Scientist to drive the development of its Vision-Language-Action foundation model for autonomous driving. You will design and train large-scale multi-modal architectures that fuse vision, language, and control across fleet data and simulations.

You will collaborate with top researchers and engineers to scale training on thousands of GPUs, develop cross-modal alignment methods, and optimize deployment while pursuing state-of-the-art results

Qualifications

  • Master’s degree or higher in Computer Science, Electrical/Computer Engineering, or related field, with 3+ years of deep learning research or productization experience.
  • Strong proficiency in PyTorch and modern transformer-based model design.
  • Experience in large-scale pretraining or multi-modal modeling (vision, language, or planning).
  • Deep understanding of representation learning, temporal modeling, and self-supervised or reinforcement learning techniques.

Responsibilities

  • Design and implement large-scale multi-modal architectures (e.g., vision-language-action transformers) for end-to-end autonomous driving.
  • Develop pretraining and fine-tuning strategies leveraging massive labeled and unlabeled fleet data (images, video, LiDAR, CAN bus, maps, human driving behaviors, etc.).
  • Research and integrate cross-modal alignment (e.g., visual grounding, temporal reasoning, policy distillation, imitation and reinforcement learning) to improve model interpretability and action quality.
  • Collaborate with infrastructure engineers to scale training across thousands of GPUs using distributed training frameworks (FSDP, DDP, etc.).
  • Conduct systematic ablation, evaluation, and visualization of model behavior across perception, reasoning, and planning tasks.
  • Contribute to model deployment optimization, including quantization, export, and latency-accuracy trade-offs for onboard execution.

Skills

PyTorch
transformer models
multi-modal modeling
distributed training
vision-language-action
reinforcement learning

Education

Master’s degree or higher in CS/EE
PhD in CS/CE/EE (preferred)

Tools

FSDP
DDP

Job description

XPENG is seeking a Machine Learning Engineer / Research Scientist to drive the development of its Vision-Language-Action foundation model for autonomous driving. You will design and train large-scale multi-modal architectures that fuse vision, language, and control across fleet data and simulations.

You will collaborate with top researchers and engineers to scale training on thousands of GPUs, develop cross-modal alignment methods, and optimize deployment while pursuing state-of-the-art results

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Engineer: Vision-Language-Action for Autonomous Driving
Staff ML Engineer: Vision-Language-Action for Autonomous Driving

XPENG • Santa Clara (CA)

On-site
USD 215,280 - 364,320
Competitive compensation package
Snacks, lunches, dinners, and fun activities
Access to massive real-world data and industry-scale compute
Staff Machine Learning Engineer - Foundation Model
Staff Machine Learning Engineer - Foundation Model

XPENG • Santa Clara (CA)

On-site
USD 215,280 - 364,320
Competitive compensation package
Snacks, lunches, dinners, and fun activities
Access to massive real-world data and industry-scale compute
Senior Machine Learning Engineer - Foundation Model
Senior Machine Learning Engineer - Foundation Model

XPENG • Santa Clara (CA)

On-site
USD 175,000 - 296,000
Massive real-world data
Industry-scale compute
Top-tier researchers and engineers
+2
Staff ML Engineer - Traffic Sign Detection for Autonomy
Staff ML Engineer - Traffic Sign Detection for Autonomy

XPENG • Santa Clara (CA)

On-site
USD 215,280 - 364,320
Snacks, lunches, dinners, and fun activities
Senior ML Engineer – World Models & Multimodal AI
Senior ML Engineer – World Models & Multimodal AI

Socket.dev • Santa Clara (CA)

On-site
USD 175,000 - 296,000
Competitive compensation
Snacks, lunches, dinners
Career growth opportunities
Senior Vision-Language-Action Engineer, Autonomous Driving
Senior Vision-Language-Action Engineer, Autonomous Driving

Socket.dev • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Catered lunch
Unlimited snacks and beverages
401(k) plan
Vision-Language-Action AI Engineer for Autonomous Driving
Vision-Language-Action AI Engineer for Autonomous Driving

Plusai • Santa Clara (CA)

On-site
USD 170,000 - 260,000
Catered free lunch
Unlimited snacks and beverages
401(k) plan
ML Engineer: Vision-Language & Motion for Autonomy
ML Engineer: Vision-Language & Motion for Autonomy

Praxis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 240,000
Senior ML Engineer, Autonomous Driving Foundation Models
Senior ML Engineer, Autonomous Driving Foundation Models

XPENG • Santa Clara (CA)

On-site
USD 244,140 - 413,160
Competitive compensation package
Snacks, lunches, dinners, and fun activities
Supportive and engaging environment
+2
ML Engineer – Vision-Language for Motion & Autonomy
ML Engineer – Vision-Language for Motion & Autonomy

Nomadic AI • San Francisco (CA)

On-site
USD 170,000 - 250,000