ML Infra Engineer - GPU Clusters & AIOps

NVIDIA Corporation

Santa Clara, Northern (CA, KY)

Hybrid

USD 124,000 - 196,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NVIDIA is seeking a Software Engineer to accelerate ML infrastructure for GPU clusters, enabling researchers to train, fine-tune, and deploy models with reduced operational overhead.

You will design, develop, and maintain scalable engineering solutions, work on Linux environments, and apply container technologies like Docker and Kubernetes. On-call support for production systems is part of the role.

Qualifications

  • BS/MS in Computer Science, Engineering, or equivalent experience.
  • 2+ years in software/platform engineering, including 1 year in ML infrastructure or distributed systems.
  • Experience in software development lifecycle on Linux-based platforms.
  • Strong coding skills in languages such as Python, C++ or Rust.
  • Experience with Docker, Kubernetes, GitLab CI, automated deployments.
  • Experience with AIOps or Agentic AI and apply it successfully in production environment.

Responsibilities

  • Understand pain points of validating, monitoring and operating GPU clusters at scale.
  • Design, develop and maintain engineering solutions to solve those pain points systematically.
  • Research in traditional AIOps and the emerging Agentic AI, and leverage it to further reduce the operation toil.
  • Participate in on-call support for systems, platforms built and owned by the team.

Skills

CS/Engineering knowledge
Python
C++
Rust
Linux

Education

BS/MS in Computer Science/Engineering

Tools

Docker
Kubernetes
GitLab CI
Slurm

Job description

NVIDIA is seeking a Software Engineer to accelerate ML infrastructure for GPU clusters, enabling researchers to train, fine-tune, and deploy models with reduced operational overhead.

You will design, develop, and maintain scalable engineering solutions, work on Linux environments, and apply container technologies like Docker and Kubernetes. On-call support for production systems is part of the role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer: GPU Clusters & Self-Serve Scale
ML Infra Engineer: GPU Clusters & Self-Serve Scale

NVIDIA • Austin (TX)

On-site
USD 124,000 - 196,000
Equity
Benefits
GPU ML Infra Engineer for Clusters & AIOps Equity
GPU ML Infra Engineer for Clusters & AIOps Equity

Socket.dev • North Carolina

On-site
USD 124,000 - 196,000
Equity
Benefits
AI Research Clusters Engineer — GPU ML Infra & AIOps
AI Research Clusters Engineer — GPU ML Infra & AIOps

NVIDIA AI • Durham (NC)

On-site
USD 120,000 - 180,000
Equity
Benefits
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

Cursor • California (MO)

On-site
USD 140,000 - 185,000
ML Infra Engineer — GPU Clusters & Distributed Systems
ML Infra Engineer — GPU Clusters & Distributed Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Industry-leading compensation and/or:?
Unlimited PTO
Top-tier medical, dental, and vision
+1
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
ML Engineer - Scalable AI Systems & GPU Orchestration
ML Engineer - Scalable AI Systems & GPU Orchestration

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Software Engineer: ML Infra
Software Engineer: ML Infra

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 120,000 - 160,000
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

cursor • New York (NY), San Francisco (CA)

On-site
USD 120,000 - 150,000