Senior AI Infra Engineer — Large-Scale GPU Cloud Equity

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NVIDIA Corporation is seeking a Senior Software Engineer to lead bring-up, benchmarking, and optimization of distributed AI training and inference workloads across GPU platforms at scale.

You will set technical direction across communication libraries, model frameworks, and inference/training stacks, perform deep performance analyses, and mentor engineers to raise the bar for the broader performance and infrastructure teams.

Qualifications

  • Bachelor’s or Master’s in Computer Science or related field
  • 8+ years of experience developing software infrastructure for large‑scale AI or HPC systems
  • Expertise debugging and triaging AI applications across the full stack—from application to hardware
  • Deep hands‑on experience with NCCL, CUDA‑aware distributed execution, and debugging multi‑GPU/multi‑node workloads
  • Proven track record of architecting, debugging, and scaling large‑scale distributed systems
  • Expert‑level Python and C/C++ programming skills
  • Experience operating workloads in scheduled, containerized cluster environments
  • Excellent analytical, debugging, and communication skills

Responsibilities

  • Lead bring‑up, validation, and debugging of large‑scale AI clusters, infrastructure, and end‑to‑end workloads
  • Benchmark and tune AI pre-training, post‑training, and inference workloads
  • Profile and optimize end‑to‑end workload performance across compute, memory, networking, and communication layers
  • Analyze scaling efficiency for distributed LLM workloads and translate findings into concrete tuning guidance
  • Own root‑cause analysis of complex failures in large distributed environments
  • Define and build the resilience and failure‑attribution stack for datacenter‑scale infrastructure
  • Build repeatable benchmark suites, automation, acceptance criteria, and qualification workflows
  • Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams
  • Deliver actionable, data‑driven recommendations based on profiling, benchmark results, and cluster characterization
  • Mentor engineers, drive technical standards, and act as a force multiplier across the broader performance and infrastructure organization

Skills

Distributed systems
Python
C/C++
Performance profiling
Debugging complex systems
Linux/Cluster environments
Communication

Education

Bachelor's or Master’s in Computer Science

Tools

NCCL
CUDA
Nsight
Docker
Kubernetes

Job description

NVIDIA Corporation is seeking a Senior Software Engineer to lead bring-up, benchmarking, and optimization of distributed AI training and inference workloads across GPU platforms at scale.

You will set technical direction across communication libraries, model frameworks, and inference/training stacks, perform deep performance analyses, and mentor engineers to raise the bar for the broader performance and infrastructure teams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infra Architect GPU & Cloud
Senior AI Infra Architect GPU & Cloud

Nvidia Corporation in • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Performance Engineer - Remote GPU Systems
Senior AI Performance Engineer - Remote GPU Systems

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 258,750 - 431,250
Equity
Comprehensive benefits
Senior AI Infra Engineer — Equity Eligible, Scale Telemetry
Senior AI Infra Engineer — Equity Eligible, Scale Telemetry

NVIDIA • California (MO)

On-site
USD 184,000 - 357,000
Senior AI Systems Performance Engineer
Senior AI Systems Performance Engineer

NVIDIA • Austin (TX)

On-site
USD 272,000 - 431,250
Equity
Benefits
Senior Systems Software Engineer: AI Infra & Kubernetes
Senior Systems Software Engineer: AI Infra & Kubernetes

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Health benefits
Flexible work arrangement
+1
Senior AI Infra Engineer | Equity Eligible, Observability Focus
Senior AI Infra Engineer | Equity Eligible, Observability Focus

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Data Storage Engineer - Cloud-Native, Equity
Senior AI Data Storage Engineer - Cloud-Native, Equity

NVIDIA • North Carolina

On-site
USD 152,000 - 241,500
Equity compensation
Benefits
Senior AI Infra Engineer — Build Telemetry & Ops, Equity
Senior AI Infra Engineer — Build Telemetry & Ops, Equity

NVIDIA • Austin (TX)

On-site
USD 184,000 - 357,000
Senior AI Performance & Efficiency Engineer - Equity Eligible
Senior AI Performance & Efficiency Engineer - Equity Eligible

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
Lead Architect, Scaled AI Inference
Lead Architect, Scaled AI Inference

Nvidia Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits package