ML Infra Engineer: GPU Clusters & Self-Serve Scale

NVIDIA

Austin (TX)

On-site

USD 124,000 - 196,000

Full time

3 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Software Engineer to accelerate ML innovation by delivering reliable GPU cluster solutions for internal researchers. You will reduce operational toil, enable self-service improvements, and empower scientists to train and deploy advanced ML models on high-performance GPU systems.

Responsibilities include designing scalable engineering solutions, researching AIOps/Agentic AI, and participating in on-call support for owned systems and platforms.

Qualifications

  • BS/MS in Computer Science, Engineering, or equivalent.
  • 2+ years in software/platform engineering, including 1 year in ML infrastructure or distributed systems.
  • Experience on Linux-based platforms and CI/CD pipelines.
  • Strong coding skills in Python, C++, or Rust.
  • Experience with Docker, Kubernetes, GitLab CI, and automated deployments.
  • Experience with AIOps or Agentic AI in production.

Responsibilities

  • Design, develop and maintain solutions to validate, monitor and operate GPU clusters at scale.
  • Research AIOps and Agentic AI to reduce operational toil.
  • Participate in on-call support for systems and platforms.

Skills

Python
C++
Rust
Docker
Kubernetes
GitLab CI
AIOps
Agentic AI
Linux

Education

BS/MS in Computer Science, Engineering, or equivalent

Tools

Slurm
Scheduling frameworks
REST API

Job description

NVIDIA is seeking a Software Engineer to accelerate ML innovation by delivering reliable GPU cluster solutions for internal researchers. You will reduce operational toil, enable self-service improvements, and empower scientists to train and deploy advanced ML models on high-performance GPU systems.

Responsibilities include designing scalable engineering solutions, researching AIOps/Agentic AI, and participating in on-call support for owned systems and platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer - GPU Clusters & AIOps
ML Infra Engineer - GPU Clusters & AIOps

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 124,000 - 196,000
GPU ML Infra Engineer for Clusters & AIOps Equity
GPU ML Infra Engineer for Clusters & AIOps Equity

Socket.dev • North Carolina

On-site
USD 124,000 - 196,000
Equity
Benefits
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

Cursor • California (MO)

On-site
USD 140,000 - 185,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
AI Research Clusters Engineer — GPU ML Infra & AIOps
AI Research Clusters Engineer — GPU ML Infra & AIOps

NVIDIA AI • Durham (NC)

On-site
USD 120,000 - 180,000
Equity
Benefits
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 150,000
Remote AI Systems Engineer — Scale GPU ML Infra
Remote AI Systems Engineer — Scale GPU ML Infra

Bright Vision Technologies • Palo Alto (CA)

On-site
USD 90,000 - 100,000
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
ML Engineer - Scalable AI Systems & GPU Orchestration
ML Engineer - Scalable AI Systems & GPU Orchestration

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits