Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA

California (MO)

On-site

USD 124,000 - 196,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity eligibility

Job summary

NVIDIA is seeking a Senior Inference Performance Engineer to push performance limits for large-scale AI inference. You will optimize AI model execution, distill insights into reusable methods, and collaborate across teams to deliver measurable improvements on NVIDIA GPUs.

You will profile using Nsight tools, explore batching, quantization, and MoE serving, and drive downstream patches that advance the public Pareto frontier while ensuring strict correctness.

Qualifications

  • Advanced degree or equivalent in CS/CE/EE or related field.
  • 3+ years of relevant engineering experience.
  • Proven knowledge of AI model execution optimization, batching, MoE, quantization.
  • Experience benchmarking GPU workloads with Nsight tools.
  • Strong Python and C++/CUDA skills.
  • Ability to communicate performance tradeoffs clearly.

Responsibilities

  • Distill performance instincts into reusable workflows and configurations for autonomous AI agents.
  • Improve throughput-latency tradeoffs in AI inference workloads and explore batching, parallelism, and quantization options.
  • Measure and optimize serving architectures across TensorRT-LLM, vLLM, and Dynamo on NVIDIA GPUs.
  • Profile workloads with Nsight Systems/Compute and interpret kernel-level data to drive fixes.
  • Land upstream improvements such as patches and deployment recipes while maintaining model correctness.

Skills

Python
C++/CUDA
Experiment design
Profiling
Technical communication

Education

BS/MS/PhD in CS/CE/EE or related

Tools

Nsight Systems
Nsight Compute
CUPTI
PyTorch profiler

Job description

NVIDIA is seeking a Senior Inference Performance Engineer to push performance limits for large-scale AI inference. You will optimize AI model execution, distill insights into reusable methods, and collaborate across teams to deliver measurable improvements on NVIDIA GPUs.

You will profile using Nsight tools, explore batching, quantization, and MoE serving, and drive downstream patches that advance the public Pareto frontier while ensuring strict correctness.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer: AI GPU Optimization&Equity
Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior GPU Systems Engineer: Scale AI Performance
Senior GPU Systems Engineer: Scale AI Performance

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Performance Tools Architect
Senior AI Performance Tools Architect

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Senior AI Performance Profiling Engineer - Equity & Benefits
Senior AI Performance Profiling Engineer - Equity & Benefits

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Inference Performance Engineer, Agent Driven Inference Optimization
Inference Performance Engineer, Agent Driven Inference Optimization

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Senior Performance Engineer: AI Workload Optimization
Senior Performance Engineer: AI Workload Optimization

NVIDIA • Redmond (WA)

On-site
USD 224,000 - 432,000
Equity
Benefits