Senior ML Engineer: LLM Quantization & In-Vehicle Deployment

XPENG

Santa Clara (CA)

On-site

USD 175,000 - 296,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

XPENG is seeking a Senior Machine Learning Engineer to lead LLM quantization and deployment for the XPENG Turing AI chip. You will quantize and optimize models, productionize PTQ/QAT workflows, and build end-to-end inference and deployment pipelines in a fast-paced autonomous mobility environment.

The role requires strong PyTorch expertise, Python software engineering, and the ability to collaborate across research, systems, and in-vehicle teams to ensure performance and reliability.

Qualifications

  • Master in CS/CE/EE, or equivalent, with 1-3 years of industry experience. Open to new graduates.
  • Strong understanding of Transformer architectures and LLM inference.
  • Hands-on experience quantizing or deploying deep learning models in production.
  • Proficiency with PyTorch and at least one inference or compilation stack.
  • Strong Python programming and software engineering skills.
  • Ability to work effectively across research, systems, infrastructure, and product teams.
  • Excellent communication and problem-solving skills, with the ability to thrive in a fast-paced and collaborative environment.
  • Experience with weight-only, activation, KV-cache, dynamic, static, or mixed-precision quantization.
  • Experience with AWQ, GPTQ, SmoothQuant, or related methods.
  • Strong numerical analysis and systems engineering skills.
  • Experience with one or more LLM runtimes, such as TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, or custom runtimes.
  • Experience deploying LLMs on resource-constrained or heterogeneous hardware.
  • Contributions to model optimization, inference, compiler, or serving projects.
  • Publications at NeurIPS, ICML, ICLR, ACL, or related conferences.

Responsibilities

  • Develop VLA inference models, ensure numerical consistency with training models, and productionize LLM quantization methods, including PTQ, QAT, mixed-precision inference, INT8, FP4, and lower-bit techniques.
  • Develop production-quality Python code with strong testing, observability, reproducibility, and failure handling.
  • Build robust model export, calibration, benchmarking, validation, and deployment pipelines.
  • Engage early with the VLA model research team to establish performance estimates and prove model feasibility.
  • Curate evaluation datasets and establish a comprehensive metric suite to systematically benchmark VLA performance.
  • Analyze numerical errors, accuracy regressions, and performance trade-offs.
  • Develop PTQ and QAT orchestration workflows.
  • Serve as the primary interface with field-testing and simulation teams for issue triage and autonomous driving performance sign-off.
  • Collaborate with the in-vehicle software team on latency analysis and issue triage.
  • Collaborate with the training infrastructure team to develop QAT and model distillation.

Skills

Transformer architectures
LLM inference
Python programming
PyTorch
Software engineering

Education

Master in CS/CE/EE or equivalent
Open to new graduates

Tools

PyTorch
Inference/compilation stack
Python
TensorRT-LLM
vLLM
llama.cpp
ONNX Runtime
TVM

Job description

XPENG is seeking a Senior Machine Learning Engineer to lead LLM quantization and deployment for the XPENG Turing AI chip. You will quantize and optimize models, productionize PTQ/QAT workflows, and build end-to-end inference and deployment pipelines in a fast-paced autonomous mobility environment.

The role requires strong PyTorch expertise, Python software engineering, and the ability to collaborate across research, systems, and in-vehicle teams to ensure performance and reliability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff LLM Quantization & Deployment Engineer
Staff LLM Quantization & Deployment Engineer

XPENG • Santa Clara (CA)

On-site
USD 215,000 - 364,000
Competitive compensation package
Infrastructures and computational资源
Staff Machine Learning Engineer - LLM Quantization & Deployment
Staff Machine Learning Engineer - LLM Quantization & Deployment

XPENG • Santa Clara (CA)

On-site
USD 215,000 - 364,000
Competitive compensation package
Infrastructures and computational资源
Senior ML Engineer, AI Foundation for Large Models
Senior ML Engineer, AI Foundation for Large Models

XPENG • Santa Clara (CA)

On-site
USD 174,720 - 295,680
Competitive compensation package
Supportive work environment
Snacks, lunches, and fun activities
Staff ML Engineer: Edge AI Quantization Lead
Staff ML Engineer: Edge AI Quantization Lead

Socket.dev • Santa Clara (CA)

On-site
USD 161,000 - 241,000
Senior ML Engineer: World Modeling for Autonomous Driving
Senior ML Engineer: World Modeling for Autonomous Driving

XPENG Jordan • Santa Clara (CA)

On-site
USD 175,000 - 296,000
Infrastructures and computational资源
Competitive compensation
Snacks and meals
Senior ML Engineer: Build Large-Scale Foundation Models
Senior ML Engineer: Build Large-Scale Foundation Models

XPENG • Santa Clara (CA)

On-site
USD 175,000 - 296,000
Snacks
Lunches
Dinners
+1
AI Intern: Edge ML & Multimodal Model Deployment
AI Intern: Edge ML & Multimodal Model Deployment

XPENG • Santa Clara (CA)

On-site
USD 70,000 - 90,000
Competitive compensation package
Snacks, lunches, dinners, and fun activities
Supportive work environment
Senior Machine Learning Engineer - AI Foundation
Senior Machine Learning Engineer - AI Foundation

XPENG • Santa Clara (CA)

On-site
USD 174,720 - 295,680
Competitive compensation package
Supportive work environment
Snacks, lunches, and fun activities
ML Engineer: World Models & Generative AI (Equity)
ML Engineer: World Models & Generative AI (Equity)

Xpeng Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 175,000 - 296,000
Competitive compensation package
Snacks, lunches, dinners, and fun δρα
Senior ML Engineer – World Models & Multimodal AI
Senior ML Engineer – World Models & Multimodal AI

Socket.dev • Santa Clara (CA)

On-site
USD 175,000 - 296,000
Competitive compensation
Snacks, lunches, dinners
Career growth opportunities