GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation

Santa Clara (CA)

On-site

USD 152,000 - 242,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking a Senior Inference Performance Engineer to push the performance limits of large-scale AI inference benchmarks. You will optimize autonomous frameworks used by AI agents to run benchmarks, profile, and tune processes, aiming to maximize throughput per GPU while maintaining model correctness.

You will collaborate with TensorRT-LLM, vLLM, and other teams to translate profiling insights into delivered performance improvements, using Nsight systems and CUDA-based tools to drive

Qualifications

  • BS/MS/PhD in CS/CE/EE/Applied Math or related field, or equivalent experience.

Responsibilities

  • Distill performance instincts into reusable workflows for autonomous AI agents.

Skills

Python
C++/CUDA
Experimental methodology
Communication
GPU architectures

Education

BS/MS/PhD in CS/CE/EE/Applied Math or related

Tools

Nsight Systems
Nsight Compute
CUPTI
PyTorch profiler

Job description

NVIDIA is seeking a Senior Inference Performance Engineer to push the performance limits of large-scale AI inference benchmarks. You will optimize autonomous frameworks used by AI agents to run benchmarks, profile, and tune processes, aiming to maximize throughput per GPU while maintaining model correctness.

You will collaborate with TensorRT-LLM, vLLM, and other teams to translate profiling insights into delivered performance improvements, using Nsight systems and CUDA-based tools to drive

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer: AI GPU Optimization&Equity
Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Inference Performance Engineer, Agent Driven Inference Optimization
Inference Performance Engineer, Agent Driven Inference Optimization

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Inference Performance Engineer, Agent Driven Inference Optimization
Inference Performance Engineer, Agent Driven Inference Optimization

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Inference Performance Engineer, AI Inference Configuration Optimization
Inference Performance Engineer, AI Inference Configuration Optimization

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior AI Performance Profiling Engineer - Equity & Benefits
Senior AI Performance Profiling Engineer - Equity & Benefits

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Inference Performance Engineer, AI Inference Configuration Optimization
Inference Performance Engineer, AI Inference Configuration Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000