Senior/Staff AI Infrastructure Engineer, Inference & Optimization

DiDi Labs

San Jose (CA)

On-site

USD 170,000 - 339,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

DiDi Labs is seeking an experienced Senior/Staff AI Infrastructure Engineer, Inference & Optimization to lead performance tuning and deployment of AI models across vehicle and cloud infrastructure. You will design high‑efficiency inference pipelines and system stability frameworks for ultra‑low latency operation in autonomous systems.

Collaborate with perception, cloud, and safety teams to accelerate model iteration, implement profiling, and optimise hardware usage on NVIDIA GPUs across Hopper

Qualifications

  • Bachelor’s or higher degree in Computer Science, Software Engineering, Systems Engineering, or a closely related field.
  • 3–8+ years of industry experience in high‑performance computing, AI infrastructure, model optimisation, or embedded deployment.
  • Strong proficiency in C++ and Python, with expertise in parallel programming (CUDA, OpenMP) and low‑level system profiling tools.
  • Deep familiarity with inference engines (TensorRT, ONNX Runtime) and specialised LLM inference/serving frameworks (vLLM, SGLang, TensorRT‑LLM).
  • Practical understanding of modern GPU hardware architectures (e.g., NVIDIA Hopper, Thor) and memory bandwidth management.
  • Demonstrated ability to diagnose complex software‑hardware integration issues and drive scalable, production‑grade solutions.

Responsibilities

  • Own the deployment, optimisation, and resource scheduling of vehicle‑side AI models, ensuring high efficiency, low latency, and robust execution within embedded constraints.
  • Lead vehicle‑side system stability initiatives, conducting independent root‑cause analysis and driving resolution for complex, system‑level performance bottlenecks and runtime anomalies.
  • Architect and scale service‑oriented deployment environments for Large Language Models (LLMs) and foundational models to support offline simulation, automated annotation, and rapid model validation.
  • Track and evaluate cutting‑edge industry methodologies, continuously integrating advanced optimisation toolchains, quantisation techniques, and execution engines.
  • Establish system‑level profiling and telemetry frameworks using CUDA tools to monitor, analyse, and maximise hardware utilisation across target GPU architectures.
  • Collaborate cross‑functionally with Autonomous Driving Perception/Prediction, Cloud Infrastructure, and Safety teams to enable rapid algorithm iteration and scalable vehicle deployment.

Skills

C++
Python
CUDA
OpenMP
GPU profiling
Root-cause analysis

Education

Bachelor's degree or higher in CS/Engineering

Tools

TensorRT
ONNX Runtime
vLLM
SGLang
TensorRT-LLM
CUDA tools

Job description

About the Company

DiDi’s autonomous driving unit was established in 2016 with the mission of developing Level 4 autonomous driving (AD) technology to make transportation safer and more efficient. In August 2019, the unit became an independent company, DiDi Autonomous Driving, dedicated to advanced AD R&D, product application, and business expansion. We believe integrating AD technology into a shared‑mobility fleet will generate immense social value. By leveraging DiDi’s specialized technology, operational expertise, and integrated ecosystem, we are positioned to build and operate a highly efficient, user‑oriented autonomous fleet.

About The Role

We are seeking an experienced and mission‑driven Senior/Staff AI Infrastructure Engineer, Inference & Optimization to lead the performance tuning, deployment, and resource scheduling of cutting‑edge AI models across on‑vehicle and cloud infrastructure. In this role, you will design high‑efficiency inference pipelines, build system‑level stability frameworks, and optimise hardware execution to ensure ultra‑low latency and rock‑solid operational reliability. You will act as a technical leader in AI infrastructure, accelerating model iteration and bridging the gap between frontier deep learning algorithms and real‑time autonomous systems.

Responsibilities
  • Own the deployment, optimisation, and resource scheduling of vehicle‑side AI models, ensuring high efficiency, low latency, and robust execution within embedded constraints.

  • Lead vehicle‑side system stability initiatives, conducting independent root‑cause analysis and driving resolution for complex, system‑level performance bottlenecks and runtime anomalies.

  • Architect and scale service‑oriented deployment environments for Large Language Models (LLMs) and foundational models to support offline simulation, automated annotation, and rapid model validation.

  • Track and evaluate cutting‑edge industry methodologies, continuously integrating advanced optimisation toolchains, quantisation techniques, and execution engines.

  • Establish system‑level profiling and telemetry frameworks using CUDA tools to monitor, analyse, and maximise hardware utilisation across target GPU architectures.

  • Collaborate cross‑functionally with Autonomous Driving Perception/Prediction, Cloud Infrastructure, and Safety teams to enable rapid algorithm iteration and scalable vehicle deployment.

Qualifications
  • Bachelor’s or higher degree in Computer Science, Software Engineering, Systems Engineering, or a closely related technical field.

  • 3–8+ years of industry experience in high‑performance computing, AI infrastructure, model optimisation, or embedded deployment.

  • Strong proficiency in C++ and Python, with solid expertise in parallel programming (CUDA, OpenMP) and low‑level system profiling tools.

  • Deep familiarity with mainstream inference engines (e.g., TensorRT, ONNX Runtime) and specialised LLM inference/serving frameworks (e.g., vLLM, SGLang, TensorRT‑LLM).

  • Practical understanding of modern GPU hardware architectures (e.g., NVIDIA Hopper, Thor) and memory bandwidth management.

  • Demonstrated ability to diagnose complex software‑hardware integration issues and drive scalable, production‑grade solutions.

Preferred Qualifications
  • Hands‑on experience optimising and deploying AI models on the NVIDIA Thor platform, including hardware resource scheduling and acceleration.

  • Proven track record of serving large foundation models (e.g., LLaMA, Qwen, GPT) in production or high‑throughput cloud pipelines using frameworks like vLLM, SGLang, TGI, or LightLLM.

  • Background in deep learning training frameworks (PyTorch) and practical experience with model quantisation (INT8/FP8/AWQ), kernel fusion, or graph compilation.

  • Experience deploying real‑time, high‑availability AI workloads in autonomous vehicles, robotics, or edge devices.

The base salary range for this full‑time position is $169,783 - $338,694 annually in addition to bonus, equity and benefits. Our salary ranges are determined by role, level, and location. Within the range, individual pay is determined by work location and additional factors, including job‑related skills, experience, and relevant education or training.

I acknowledge that prior to submitting this application, I have read and accepted the Privacy Notice for California Residents which is available on https://v.didi.cn/AQnxlBa

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

DiDi • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization

Didi Labs • San Jose (CA)

On-site
USD 170,000 - 351,000
Staff/Principal AI Transformation Engineer
Staff/Principal AI Transformation Engineer

Didi Labs • San Jose (CA)

On-site
USD 255,000 - 351,000
Principal AI Transformation Engineer
Principal AI Transformation Engineer

Worky • San Jose (CA)

On-site
USD 255,000 - 351,000
Sr. / Staff Software Engineer, Infrastructure (Autonomy)
Sr. / Staff Software Engineer, Infrastructure (Autonomy)

Didi Labs • San Jose (CA)

On-site
USD 170,000 - 339,000
Senior AI Infrastructure Engineer, Inference & Optimization
Senior AI Infrastructure Engineer, Inference & Optimization

Didi Labs • San Jose (CA)

On-site
USD 170,000 - 351,000
Senior AI Inference & Optimization Infrastructure Engineer
Senior AI Inference & Optimization Infrastructure Engineer

DiDi • San Jose (CA)

On-site
USD 170,000 - 351,000
Sr. / Staff Software Engineer, Infrastructure (Autonomy)
Sr. / Staff Software Engineer, Infrastructure (Autonomy)

DiDi • San Jose (CA)

On-site
USD 170,000 - 339,000
Software Engineer – Map Fusion & Planning
Software Engineer – Map Fusion & Planning

Socket.dev • San Jose (CA)

On-site
USD 129,189 - 214,776
Senior AI Inference & Optimization Engineer — Autonomy
Senior AI Inference & Optimization Engineer — Autonomy

DiDi Labs • San Jose (CA)

On-site
USD 170,000 - 339,000