Senior AI Infra Engineer-Distributed GPU Clusters (Equity)

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 184,000 - 356,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Gruppe in Santa Clara is seeking a Senior Software Engineer to lead the optimization of distributed training across large-scale GPU platforms. Candidates should have substantial experience in AI applications and technical leadership.

This role involves profiling end-to-end workloads, debugging complex systems, and delivering actionable insights to drive performance improvements. You will also mentor engineers and define standards that enhance engineering practices.

An inclusive work environment ensures diverse perspectives contribute to our pioneering AI solutions.

Qualifications

  • 8+ years of experience developing software for large-scale AI or HPC systems.
  • Expertise debugging AI applications from the application layer to hardware.
  • Deep hands-on experience with NCCL and multi-GPU workloads.

Responsibilities

  • Lead validation and debugging of large-scale AI clusters.
  • Benchmark AI workloads using PyTorch and NVIDIA AI software.
  • Profile workload performance using tools like Nsight Systems.

Skills

Technical leadership
AI applications debugging
Python programming
C/C++ programming
Analytical skills

Education

Bachelor's or Master's in Computer Science or related field

Tools

NVIDIA software stacks
CUDA
NCCL

Job description

NVIDIA Gruppe in Santa Clara is seeking a Senior Software Engineer to lead the optimization of distributed training across large-scale GPU platforms. Candidates should have substantial experience in AI applications and technical leadership.

This role involves profiling end-to-end workloads, debugging complex systems, and delivering actionable insights to drive performance improvements. You will also mentor engineers and define standards that enhance engineering practices.

An inclusive work environment ensures diverse perspectives contribute to our pioneering AI solutions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Performance & Efficiency Engineer - Equity Eligible
Senior AI Performance & Efficiency Engineer - Equity Eligible

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
Senior HPC-AI Systems Architect (Equity)
Senior HPC-AI Systems Architect (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Senior AI Infra Architect for Multi-GPU Training + Equity
Senior AI Infra Architect for Multi-GPU Training + Equity

NVIDIA Gruppe • California (MO)

On-site
USD 224,000 - 431,000
Equity
Benefits
Senior AI Infra Systems Engineer - Equity
Senior AI Infra Systems Engineer - Equity

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 260,000
Equity
Senior AI Compute Engineer - HPC Infra & Linux
Senior AI Compute Engineer - HPC Infra & Linux

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 148,000 - 288,000
Senior AI Infra Engineer - Large-Scale DGX Cloud (Equity)
Senior AI Infra Engineer - Large-Scale DGX Cloud (Equity)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Equity options
Comprehensive benefits package
Senior GPU AI & Quant HPC Engineer
Senior GPU AI & Quant HPC Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 287,500
Equity
Benefits
Senior GPU Performance Architect - Scale, Equity Eligible
Senior GPU Performance Architect - Scale, Equity Eligible

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Infra Engineer - Scalable Cloud Platform
Senior AI Infra Engineer - Scalable Cloud Platform

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infrastructure Architect – GPU Clusters
Senior AI Infrastructure Architect – GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Equity
Benefits