Senior AI Infrastructure Engineer - Real-Time Inference

Didi Labs

San Jose (CA)

Hybrid

USD 170,000 - 351,000

Full time

12 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

DiDi seeks a Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization in San Jose, CA, to lead performance tuning, deployment, and resource scheduling of AI models across vehicle and cloud infrastructure.

You will design high‑efficiency inference pipelines, build system stability frameworks, and optimize hardware execution for ultra‑low latency, collaborating with perception/prediction, cloud, and safety teams to enable rapid model iteration and scalable deployment.

Qualifications

  • Master’s or higher degree in Computer Science, Software Engineering, Systems Engineering, or a closely related technical field.
  • 3-8+ years of industry experience in high‑performance computing, AI infrastructure, model optimization, or embedded deployment.
  • Strong proficiency in C++ and Python, with solid expertise in parallel programming (CUDA, OpenMP) and low‑level system profiling tools.
  • Deep familiarity with mainstream inference engines (e.g., TensorRT, ONNX Runtime) and specialized LLM inference/serving frameworks (e.g., vLLM, SGLang, TensorRT‑LLM).

Responsibilities

  • Own the deployment, optimization, and resource scheduling of vehicle‑side AI models, ensuring high efficiency, low latency, and robust execution within embedded constraints.
  • Lead vehicle‑side system stability initiatives, conducting independent root‑cause analysis and driving resolution for complex, system‑level performance bottlenecks.
  • Architect and scale service‑oriented deployment environments for Large Language Models (LLMs) and foundational models to support offline simulation, automated annotation, and rapid model validation.
  • Track and evaluate cutting‑edge industry methodologies, continuously integrating advanced optimization toolchains, quantization techniques, and execution engines.
  • Establish system‑level profiling and telemetry frameworks using CUDA tools to monitor, analyze, and maximize hardware utilization across target GPU architectures.
  • Collaborate cross‑functionally with Autonomous Driving Perception/Prediction, Cloud Infrastructure, and Safety teams to enable rapid algorithm iteration and scalable vehicle deployment.

Skills

C++
Python
CUDA
OpenMP
System profiling tools

Education

Master's degree in CS/SE/related field

Tools

TensorRT
ONNX Runtime
vLLM
SGLang
TensorRT-LLM

Job description

DiDi seeks a Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization in San Jose, CA, to lead performance tuning, deployment, and resource scheduling of AI models across vehicle and cloud infrastructure.

You will design high‑efficiency inference pipelines, build system stability frameworks, and optimize hardware execution for ultra‑low latency, collaborating with perception/prediction, cloud, and safety teams to enable rapid model iteration and scalable deployment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer, Inference & Optimization
Senior AI Infrastructure Engineer, Inference & Optimization

Didi Labs • San Jose (CA)

On-site
USD 170,000 - 351,000
Lead AI Inference & Optimization Engineer (Vehicle & Edge)
Lead AI Inference & Optimization Engineer (Vehicle & Edge)

didi • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization New San Jose, CA
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization New San Jose, CA

Didi Labs • San Jose (CA)

Hybrid
USD 170,000 - 351,000
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

Didi Labs • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

didi • San Jose (CA)

On-site
USD 170,000 - 351,000
AI Transformation Architect — Enterprise-Scale Platform
AI Transformation Architect — Enterprise-Scale Platform

Didi Labs • San Jose (CA)

Hybrid
USD 255,000 - 351,000
AI Infrastructure Performance & Modeling Lead
AI Infrastructure Performance & Modeling Lead

Renice AI • Mountain View (CA)

Hybrid
USD 180,000 - 280,000
AI Engineer: Model Training, Inference & GPU Infra
AI Engineer: Model Training, Inference & GPU Infra

Agentrys • San Jose (CA), Northern (KY)

Hybrid
USD 170,000 - 210,000
AI Inference Infrastructure Engineer
AI Inference Infrastructure Engineer

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Senior AI Deployment Leader: Real-Time Inference
Senior AI Deployment Leader: Real-Time Inference

General Motors • Sunnyvale (CA), Northern (KY)

Hybrid
USD 296,000 - 454,000