Senior AI Inference Optimization Engineer

Nvidia Corporation in

Santa Clara (CA)

Hybrid

USD 124,000 - 196,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits package
Hybrid work model

Job summary

NVIDIA is seeking a Senior Inference Performance Engineer to push the performance limits of AI inference on NVIDIA GPUs. This role focuses on autonomous optimization workflows, benchmarking, and tuning configurations using an optimization framework across TensorRT-LLM, vLLM, and related systems.

The position requires strong Python and C++/CUDA skills, extensive GPU profiling experience with Nsight tools, and a rigorous experimental methodology.

Qualifications

  • BS/MS/PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Math, or related field, or equivalent experience.
  • 3+ years of relevant engineering experience.
  • Strong Python engineering skills and ability to modify large C++/CUDA codebases.
  • Hands-on benchmarking and profiling GPU workloads with Nsight Systems/Compute, CUPTI, or PyTorch profiler.
  • Excellent written and verbal communication skills to explain performance tradeoffs.

Responsibilities

  • Distill performance instincts into reusable, autonomous workflows for AI agents to run benchmarks and tune configurations.
  • Improve AI inference workloads to increase throughput per GPU and user interactivity through batching, KV cache, quantization, and spec decoding options.
  • Measure and optimize serving architectures across TensorRT-LLM, SGLang, vLLM, and Dynamo on NVIDIA GPUs.
  • Profile workloads using Nsight tools and kernel traces; apply roofline/speed-of-light analysis to drive measured gains.
  • Collaborate with TensorRT-LLM, SGLang, vLLM, kernel, benchmarking, and GPU architecture teams to deliver improvements.

Job description

NVIDIA is seeking a Senior Inference Performance Engineer to push the performance limits of AI inference on NVIDIA GPUs. This role focuses on autonomous optimization workflows, benchmarking, and tuning configurations using an optimization framework across TensorRT-LLM, vLLM, and related systems.

The position requires strong Python and C++/CUDA skills, extensive GPU profiling experience with Nsight tools, and a rigorous experimental methodology.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer: AI GPU Optimization&Equity
Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Inference Performance Engineer, AI Inference Configuration Optimization
Inference Performance Engineer, AI Inference Configuration Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Inference Performance Engineer, Agent Driven Inference Optimization
Inference Performance Engineer, Agent Driven Inference Optimization

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Inference Performance Engineer, AI Inference Configuration Optimization
Inference Performance Engineer, AI Inference Configuration Optimization

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
AI Inference Performance Engineer - New College Grad 2026
AI Inference Performance Engineer - New College Grad 2026

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Senior AI Inference Performance Product Manager
Senior AI Inference Performance Product Manager

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 208,000 - 328,000