Senior ML Engineer: LLM Inference & Quantization

XPENG

Santa Clara (CA)

On-site

USD 175,000 - 296,000

Full time

13 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Snacks and meals
Cutting-edge technologies
Top talents collaboration

Job summary

XPENG, a leading smart technology company, seeks a qualified engineer to advance VLA inference and model quantization for next-gen XPENG Turing AI chip. You will develop production-grade Python code, build end-to-end pipelines, and work with research, systems, and in-vehicle teams.

The role focuses on PTQ and QAT workflows, performance benchmarking, and validating models in real-world simulations. Strong PyTorch experience, quantization know-how, and collaboration skills are essential for

Qualifications

  • Master in CS/CE/EE, or equivalent, with 1-3 years of industry experience.
  • Strong understanding of Transformer architectures and LLM inference.
  • Hands-on experience quantizing or deploying deep learning models in production.
  • Proficiency with PyTorch and at least one inference or compilation stack.
  • Strong Python programming and software engineering skills.
  • Ability to work effectively across research, systems, infrastructure, and product teams.
  • Excellent communication and problem-solving skills, with the ability to thrive in a fast-paced and collaborative environment.

Responsibilities

  • Develop VLA inference models, ensure numerical consistency with training models, and productionize LLM quantization methods.
  • Develop production-quality Python code with strong testing, observability, reproducibility, and failure handling.
  • Build robust model export, calibration, benchmarking, validation, and deployment pipelines.
  • Engage early with the VLA model research team to establish performance estimates and prove model feasibility.
  • Curate evaluation datasets and establish a comprehensive metric suite to benchmark performance.
  • Analyze numerical errors, accuracy regressions, and performance trade-offs.
  • Develop PTQ and QAT orchestration workflows.
  • Serve as the primary interface with field-testing and simulation teams for issue triage and autonomous driving performance sign-off.
  • Collaborate with the in-vehicle software team on latency analysis and issue triage.
  • Collaborate with the training infrastructure team to develop QAT and model distillation.

Skills

Transformer architectures
LLM inference
Python programming
PyTorch
Software engineering
Communication
Cross-functional collaboration

Education

Master's degree in CS/CE/EE

Tools

TensorRT-LLM
vLLM
llama.cpp
ONNX Runtime

Job description

XPENG, a leading smart technology company, seeks a qualified engineer to advance VLA inference and model quantization for next-gen XPENG Turing AI chip. You will develop production-grade Python code, build end-to-end pipelines, and work with research, systems, and in-vehicle teams.

The role focuses on PTQ and QAT workflows, performance benchmarking, and validating models in real-world simulations. Strong PyTorch experience, quantization know-how, and collaboration skills are essential for

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer - LLM Quantization & Deployment
Senior Machine Learning Engineer - LLM Quantization & Deployment

XPENG • Santa Clara (CA)

On-site
USD 175,000 - 296,000
Competitive compensation
Snacks and meals
Cutting-edge technologies
+1
Principal AI Systems Engineer - LLM Inference (Cloud)
Principal AI Systems Engineer - LLM Inference (Cloud)

Qualcomm • San Diego (CA)

On-site
USD 224,000 - 335,000
Applied AI Engineer: LLMs, Retrieval & MLOps
Applied AI Engineer: LLMs, Retrieval & MLOps

Quantum Technologies. LLC • Austin (TX), Northern (KY)

Hybrid
USD 120,000 - 160,000
Equal opportunity employer
Senior Applied Scientist — Optical AI Quantization
Senior Applied Scientist — Optical AI Quantization

Neurophos • Austin (TX)

On-site
USD 180,000 - 260,000
Health plan premiums coverage
Unlimited PTO
401(k) matching
+2
Senior NLP & AI/ML Engineer
Senior NLP & AI/ML Engineer

Quantum Technologies. LLC • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 190,000
Staff AI Model Optimization Architect for LLMs & Multimodal
Staff AI Model Optimization Architect for LLMs & Multimodal

Qualcomm • Austin (TX)

On-site
USD 158,000 - 238,000
AI Foundation ML Engineer – Massive Training & Inference
AI Foundation ML Engineer – Massive Training & Inference

XPENG • Santa Clara (CA)

On-site
USD 175,000 - 296,000
Snacks
Lunches
Dinners
+1
Senior AI Inference Architect for Production LLMs
Senior AI Inference Architect for Production LLMs

Crusoe • San Francisco (CA)

On-site
USD 250,000 - 300,000
Competitive compensation and equity
Restricted Stock Units
Paid time off & holidays
+13
Research Scientist: Efficient AI Inference
Research Scientist: Efficient AI Inference

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 140,000 - 220,000
AI Infra Performance Intern: Embedded & HPC Optimizations
AI Infra Performance Intern: Embedded & HPC Optimizations

XPENG • Santa Clara (CA)

On-site
USD 100,000 - 150,000
Supportive environment
Computational resources
Opportunities for growth
+2