Senior AI Inference & Optimization Engineer — Autonomy

DiDi Labs

San Jose (CA)

On-site

USD 170,000 - 339,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

DiDi Labs is seeking an experienced Senior/Staff AI Infrastructure Engineer, Inference & Optimization to lead performance tuning and deployment of AI models across vehicle and cloud infrastructure. You will design high‑efficiency inference pipelines and system stability frameworks for ultra‑low latency operation in autonomous systems.

Collaborate with perception, cloud, and safety teams to accelerate model iteration, implement profiling, and optimise hardware usage on NVIDIA GPUs across Hopper

Qualifications

  • Bachelor’s or higher degree in Computer Science, Software Engineering, Systems Engineering, or a closely related field.
  • 3–8+ years of industry experience in high‑performance computing, AI infrastructure, model optimisation, or embedded deployment.
  • Strong proficiency in C++ and Python, with expertise in parallel programming (CUDA, OpenMP) and low‑level system profiling tools.
  • Deep familiarity with inference engines (TensorRT, ONNX Runtime) and specialised LLM inference/serving frameworks (vLLM, SGLang, TensorRT‑LLM).
  • Practical understanding of modern GPU hardware architectures (e.g., NVIDIA Hopper, Thor) and memory bandwidth management.
  • Demonstrated ability to diagnose complex software‑hardware integration issues and drive scalable, production‑grade solutions.

Responsibilities

  • Own the deployment, optimisation, and resource scheduling of vehicle‑side AI models, ensuring high efficiency, low latency, and robust execution within embedded constraints.
  • Lead vehicle‑side system stability initiatives, conducting independent root‑cause analysis and driving resolution for complex, system‑level performance bottlenecks and runtime anomalies.
  • Architect and scale service‑oriented deployment environments for Large Language Models (LLMs) and foundational models to support offline simulation, automated annotation, and rapid model validation.
  • Track and evaluate cutting‑edge industry methodologies, continuously integrating advanced optimisation toolchains, quantisation techniques, and execution engines.
  • Establish system‑level profiling and telemetry frameworks using CUDA tools to monitor, analyse, and maximise hardware utilisation across target GPU architectures.
  • Collaborate cross‑functionally with Autonomous Driving Perception/Prediction, Cloud Infrastructure, and Safety teams to enable rapid algorithm iteration and scalable vehicle deployment.

Skills

C++
Python
CUDA
OpenMP
GPU profiling
Root-cause analysis

Education

Bachelor's degree or higher in CS/Engineering

Tools

TensorRT
ONNX Runtime
vLLM
SGLang
TensorRT-LLM
CUDA tools

Job description

DiDi Labs is seeking an experienced Senior/Staff AI Infrastructure Engineer, Inference & Optimization to lead performance tuning and deployment of AI models across vehicle and cloud infrastructure. You will design high‑efficiency inference pipelines and system stability frameworks for ultra‑low latency operation in autonomous systems.

Collaborate with perception, cloud, and safety teams to accelerate model iteration, implement profiling, and optimise hardware usage on NVIDIA GPUs across Hopper

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference & Optimization Infrastructure Engineer
Senior AI Inference & Optimization Infrastructure Engineer

DiDi • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior AI Infrastructure Engineer, Inference & Optimization
Senior AI Infrastructure Engineer, Inference & Optimization

Didi Labs • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

DiDi • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

Didi Labs • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior/Staff AI Infrastructure Engineer, Inference & Optimization
Senior/Staff AI Infrastructure Engineer, Inference & Optimization

DiDi Labs • San Jose (CA)

On-site
USD 170,000 - 339,000
Senior AI/Systems Engineer — Autonomous Driving (Equity)
Senior AI/Systems Engineer — Autonomous Driving (Equity)

NVIDIA AI • Santa Clara (CA)

On-site
USD 240,000 - 340,000
Equity
Benefits
Remote Senior AI Inference Optimization Engineer
Remote Senior AI Inference Optimization Engineer

DigitalOcean • San Francisco (CA)

On-site
USD 191,000 - 239,000
Equity compensation
Remote work
Senior AI Systems Engineer: Automotive Inference
Senior AI Systems Engineer: Automotive Inference

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Staff Engineer, AI Inference Optimization
Staff Engineer, AI Inference Optimization

DigitalOcean • Boston (MA)

On-site
USD 191,000 - 239,000
Equity compensation
Bonus potential
Conference reimbursement
+2
Senior Infra Software Architect – Autonomy Platform
Senior Infra Software Architect – Autonomy Platform

DiDi • San Jose (CA)

On-site
USD 170,000 - 339,000