Senior AI Infrastructure Engineer, Inference & Optimization

Didi Labs

San Jose (CA)

On-site

USD 170,000 - 351,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

DiDi Autonomous Driving is seeking a Senior/Sr. Staff AI Infrastructure Engineer to lead AI model deployment, optimization, and resource scheduling across on-vehicle and cloud infrastructure.

You will design high-efficiency inference pipelines and robust, low-latency systems for embedded constraints. The role requires a Master’s degree and 3–8+ years in high-performance AI infra, with expertise in C++, Python, CUDA, and modern inference engines.

Qualifications

  • Master’s or higher in CS/Software/Systems Engineering or related field.
  • 3–8+ years in high-performance AI infrastructure, model optimization, or embedded deployment.
  • Strong proficiency in C++ and Python, CUDA/OpenMP, and low-level profiling tools.

Responsibilities

  • Own deployment, optimization, and resource scheduling of vehicle-side AI models within embedded constraints.
  • Lead vehicle-side system stability initiatives; diagnose and resolve complex bottlenecks and runtime anomalies.
  • Architect and scale deployment environments for LLMs to support offline simulation, annotation, and validation.
  • Track and evaluate cutting-edge methodologies; incorporate optimization toolchains and quantization techniques.
  • Establish system-level profiling/telemetry using CUDA tools to maximize GPU utilization.
  • Collaborate with AD Perception/Prediction, Cloud Infra, and Safety teams for rapid iteration.

Skills

C++
Python
CUDA
OpenMP
Profiling
System design
Root-cause analysis
GPU architectures

Education

Master's degree in CS/Engineering

Tools

TensorRT
ONNX Runtime
vLLM
SGLang
TensorRT-LLM
PyTorch
CUDA tools

Job description

DiDi Autonomous Driving is seeking a Senior/Sr. Staff AI Infrastructure Engineer to lead AI model deployment, optimization, and resource scheduling across on-vehicle and cloud infrastructure.

You will design high-efficiency inference pipelines and robust, low-latency systems for embedded constraints. The role requires a Master’s degree and 3–8+ years in high-performance AI infra, with expertise in C++, Python, CUDA, and modern inference engines.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead AI Inference & Optimization Engineer (Vehicle & Edge)
Lead AI Inference & Optimization Engineer (Vehicle & Edge)

didi • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

didi • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

Didi Labs • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior AI/Systems Engineer — Autonomous Driving (Equity)
Senior AI/Systems Engineer — Autonomous Driving (Equity)

NVIDIA AI • Santa Clara (CA)

On-site
USD 240,000 - 340,000
Equity
Benefits
Senior AI Engineer — Autonomous Systems & Inference
Senior AI Engineer — Autonomous Systems & Inference

AMD • San Jose (CA)

On-site
USD 180,000 - 240,000
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,200 - 239,000
Equity compensation
Remote work
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

DigitalOcean • Seattle (WA)

On-site
USD 191,000 - 239,000
Edge DL Systems Engineer - Real-Time AI Inference
Edge DL Systems Engineer - Real-Time AI Inference

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Benefits
Staff Engineer, AI Inference Optimization
Staff Engineer, AI Inference Optimization

DigitalOcean • Boston (MA)

On-site
USD 191,200 - 239,000
Equity compensation
Bonus potential
Conference reimbursement
+2
Senior Principal AI Inference Engineer - Remote/Edge
Senior Principal AI Inference Engineer - Remote/Edge

Cerence Inc. • Burlington (MA)

Hybrid
USD 185,000 - 280,000
Annual bonus opportunity
Insurance coverage
Paid time off
+1