Distributed AI Inference Performance Engineer

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Gruppe in Santa Clara is looking for an engineer to optimize and benchmark GenAI inference on cutting-edge accelerators. This role involves owning end-to-end optimization pipelines and defining next-generation inference benchmarks across multiple platforms.

The ideal candidate will have a strong background in software development, particularly in Python or C++, and a deep understanding of LLM/VLM architectures. Join our team to influence the ecosystem and contribute to impactful open-source projects.

Qualifications

  • 2+ years of relevant software development experience.
  • Strong skills in Python or C++ programming and software design.
  • Deep understanding of LLM/VLM architectures and inference mechanics.

Responsibilities

  • Drive industry benchmark results and own the optimization pipeline.
  • Define next-generation inference benchmarks and shape AI use cases.
  • Design distributed inference across GPU clusters.

Skills

Python programming
C++ programming
Software design
Software engineering
Deep Learning frameworks (e.g., PyTorch, JAX)

Education

BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering

Job description

NVIDIA Gruppe in Santa Clara is looking for an engineer to optimize and benchmark GenAI inference on cutting-edge accelerators. This role involves owning end-to-end optimization pipelines and defining next-generation inference benchmarks across multiple platforms.

The ideal candidate will have a strong background in software development, particularly in Python or C++, and a deep understanding of LLM/VLM architectures. Join our team to influence the ecosystem and contribute to impactful open-source projects.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference Performance Engineer - New College Grad 2026
AI Inference Performance Engineer - New College Grad 2026

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Inference Performance Engineer: AI GPU Optimization&Equity
Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior AI Infra Engineer — Distributed Training & Inference
Senior AI Infra Engineer — Distributed Training & Inference

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Comprehensive benefits package
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
GPU Inference Engineer — Deep Learning
GPU Inference Engineer — Deep Learning

2100 NVIDIA USA • California (MO)

On-site
USD 124,000 - 242,000
Equity
Benefits
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation