Senior AI Efficiency Engineer for GPU Clusters & Research

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 152,000 - 288,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NVIDIA Corporation seeks a Senior AI/ML Performance and Efficiency Engineer for GPU Clusters to advance AI efficiency across research workloads. You will partner with researchers to detect and fix infrastructure and application bottlenecks, delivering scalable improvements on NVIDIA GPUs.

The role demands 5+ years in managing large-scale compute infra, strong ML tooling knowledge, and hands-on experience with NCCL and NSight tools.

Qualifications

  • BS or similar background in Computer Science or related area (or equivalent experience).
  • Minimum 5+ years of experience designing and operating large scale compute infrastructure.
  • Strong understanding of modern ML techniques and tools.
  • Experience investigating, and resolving, training & inference performance end to end.
  • Debugging and optimization experience with NSight Systems and NSight Compute.
  • Experience with debugging large-scale distributed training using NCCL.
  • Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with cloud computing platforms (e.g., AWS, GCP, Azure).

Responsibilities

  • Collaborate with AI/ML researchers to improve efficiency of ML models, boosting productivity and reducing costs.
  • Build tools and frameworks; apply ML techniques to detect bottlenecks and deliver improvements.
  • Work on diverse ML workloads across Robotics, Autonomous vehicles, LLMs, Videos, and more.
  • Collaborate across engineering to optimize hardware, software, and infrastructure usage.
  • Monitor fleet-wide utilization, identify inefficiencies, and deliver scalable solutions.
  • Stay updated on AI/ML developments and advocate integration into the org.

Skills

Python
Go
Bash
NCCL
NSight Compute

Education

BS in CS or related

Tools

PyTorch
TensorFlow
InfiniBand
Lustre

Job description

NVIDIA Corporation seeks a Senior AI/ML Performance and Efficiency Engineer for GPU Clusters to advance AI efficiency across research workloads. You will partner with researchers to detect and fix infrastructure and application bottlenecks, delivering scalable improvements on NVIDIA GPUs.

The role demands 5+ years in managing large-scale compute infra, strong ML tooling knowledge, and hands-on experience with NCCL and NSight tools.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Performance & Efficiency Engineer - Equity Eligible
Senior AI Performance & Efficiency Engineer - Equity Eligible

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
Senior AI Performance and Efficiency Engineer
Senior AI Performance and Efficiency Engineer

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
Senior AI Performance and Efficiency Engineer
Senior AI Performance and Efficiency Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Remote AI Research Clusters Engineer - ML Infra & GPU
Remote AI Research Clusters Engineer - ML Infra & GPU

NEPSE Trading • Northern (KY)

Hybrid
USD 124,000 - 196,000
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity and benefits
Senior AI Networking Performance Architect — Equity
Senior AI Networking Performance Architect — Equity

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits
Senior HPC Performance Engineer — AI for Science (Equity)
Senior HPC Performance Engineer — AI for Science (Equity)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Infrastructure Architect – GPU Clusters
Senior AI Infrastructure Architect – GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 356,500
Equity
Benefits
AI Research Clusters Engineer — GPU ML Infra & AIOps
AI Research Clusters Engineer — GPU ML Infra & AIOps

NVIDIA AI • Durham (NC)

On-site
USD 120,000 - 180,000
Equity
Benefits
Senior AI Systems Performance Engineer
Senior AI Systems Performance Engineer

NVIDIA • Austin (TX)

On-site
USD 272,000 - 431,250
Equity
Benefits