Lead AI Inference & Optimization Engineer (Vehicle & Edge)

didi

San Jose (CA)

On-site

USD 170,000 - 351,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

DiDi seeks a Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization to lead performance tuning, deployment, and resource scheduling of vehicle-side and cloud AI models.

You will design efficient inference pipelines and tooling to ensure ultra-low latency in autonomous driving, while collaborating with perception/prediction, cloud infra, and safety teams. Ideal candidates have 3–8+ years in AI infrastructure, strong C++/Python, CUDA, and experience with TensorRT/ONNX Runtime;

Qualifications

  • Master's degree or higher in CS, software or systems engineering or related field.
  • 3–8+ years in high-performance computing, AI infrastructure, model optimization, or embedded deployment.
  • Strong proficiency in C++ and Python, with CUDA/OpenMP and low-level profiling tools.
  • Experience with TensorRT, ONNX Runtime and LLM inference frameworks (e.g., vLLM, TensorRT-LLM).
  • Familiarity with NVIDIA Hopper/Thor architectures and memory bandwidth management.
  • Ability to diagnose complex software-hardware integration issues and deliver production-grade solutions.

Responsibilities

  • Own deployment, optimization, and resource scheduling of vehicle-side AI models within embedded constraints.
  • Lead vehicle-side system stability initiatives and drive resolution for performance bottlenecks.
  • Architect and scale service-oriented deployment environments for LLMs and foundational models.
  • Track and evaluate industry methodologies, integrating optimization toolchains and quantization.
  • Establish profiling and telemetry with CUDA tools to maximize GPU utilization across architectures.
  • Collaborate cross-functionally with Autonomous Driving Perception/Prediction, Cloud Infrastructure, and Safety teams to enable rapid algorithm iteration and scalable deployment.

Skills

C++
Python
CUDA
OpenMP
Performance tuning
Cross-functional

Education

Master's or higher in CS/Software/Systems Eng

Tools

TensorRT
ONNX Runtime
vLLM
SGLang
TensorRT-LLM

Job description

DiDi seeks a Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization to lead performance tuning, deployment, and resource scheduling of vehicle-side and cloud AI models.

You will design efficient inference pipelines and tooling to ensure ultra-low latency in autonomous driving, while collaborating with perception/prediction, cloud infra, and safety teams. Ideal candidates have 3–8+ years in AI infrastructure, strong C++/Python, CUDA, and experience with TensorRT/ONNX Runtime;

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer, Inference & Optimization
Senior AI Infrastructure Engineer, Inference & Optimization

Didi Labs • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

didi • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

Didi Labs • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior AI/Systems Engineer — Autonomous Driving (Equity)
Senior AI/Systems Engineer — Autonomous Driving (Equity)

NVIDIA AI • Santa Clara (CA)

On-site
USD 240,000 - 340,000
Equity
Benefits
Remote Senior AI Deployment Lead — In-Vehicle Inference
Remote Senior AI Deployment Lead — In-Vehicle Inference

General Motors • United States

On-site
USD 296,000 - 454,000
Remote work
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,000 - 239,000
Equity compensation
Remote work
Senior Principal AI Inference Engineer - Remote/Edge
Senior Principal AI Inference Engineer - Remote/Edge

Cerence Inc. • Burlington (MA)

Hybrid
USD 185,000 - 280,000
Annual bonus opportunity
Insurance coverage
Paid time off
+1
Senior AI Systems Engineer: Automotive Inference
Senior AI Systems Engineer: Automotive Inference

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Senior AI Deployment Leader: Real-Time Inference
Senior AI Deployment Leader: Real-Time Inference

General Motors • Sunnyvale (CA), Northern (KY)

Hybrid
USD 296,000 - 454,000
Senior DL Inference Engineer – Automotive, C++, TensorRT Equity
Senior DL Inference Engineer – Automotive, C++, TensorRT Equity

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Equity