GPU AI Research Clusters Engineer — Scale ML Infra

NVIDIA

Westford (MA)

On-site

USD 124,000 - 196,000

Full time

2 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity

Job summary

NVIDIA in Massachusetts is seeking a Software Engineer to accelerate ML innovation by delivering reliable GPU clusters for internal researchers. You will design, implement, and operate scalable systems that reduce toil and enable self-service improvements across the ML stack.

You will work with the AI Platform team to validate, monitor, and operate GPU clusters, explore AIOps and Agentic AI, and participate in on-call support.

Qualifications

  • BS/MS in Computer Science, Engineering, or equivalent experience.
  • 2+ years in software/platform engineering, incl. 1 year in ML infrastructure or distributed systems.
  • Experience in software development lifecycle on Linux-based platforms.
  • Strong coding skills in Python, C++ or Rust.
  • Experience with Docker, Kubernetes, GitLab CI, automated deployments.
  • Experience with AIOps or Agentic AI in production.

Responsibilities

  • Collaborate across the AI Platform to understand the pain points in validating, monitoring and operating GPU clusters at scale.
  • Design, develop and maintain engineering solutions to reduce operational toil.
  • Research traditional AIOps and Agentic AI to further reduce toil.
  • Participate in on-call support for systems, platforms owned by the team.

Skills

Python
C++
Rust
Linux

Education

BS/MS in Computer Science, Engineering, or equivalent experience

Tools

Docker
Kubernetes
GitLab CI

Job description

NVIDIA in Massachusetts is seeking a Software Engineer to accelerate ML innovation by delivering reliable GPU clusters for internal researchers. You will design, implement, and operate scalable systems that reduce toil and enable self-service improvements across the ML stack.

You will work with the AI Platform team to validate, monitor, and operate GPU clusters, explore AIOps and Agentic AI, and participate in on-call support.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer: GPU Clusters & Self-Serve Scale
ML Infra Engineer: GPU Clusters & Self-Serve Scale

NVIDIA • Austin (TX)

On-site
USD 124,000 - 196,000
Equity
Benefits
GPU ML Infra Engineer for Clusters & AIOps Equity
GPU ML Infra Engineer for Clusters & AIOps Equity

Socket.dev • North Carolina

On-site
USD 124,000 - 196,000
Equity
Benefits
AI Research Clusters Engineer — GPU ML Infra & AIOps
AI Research Clusters Engineer — GPU ML Infra & AIOps

NVIDIA AI • Durham (NC)

On-site
USD 120,000 - 180,000
Equity
Benefits
ML Infra Engineer - GPU Clusters & AIOps
ML Infra Engineer - GPU Clusters & AIOps

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 124,000 - 196,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
AI/ML Infra Engineer — GPU Clusters & HPC
AI/ML Infra Engineer — GPU Clusters & HPC

NVIDIA AI • Redmond (WA)

On-site
USD 120,000 - 180,000
Competitive salaries
Comprehensive benefits package
Equity
Remote AI Systems Engineer — Scale GPU ML Infra
Remote AI Systems Engineer — Scale GPU ML Infra

Bright Vision Technologies • Palo Alto (CA)

On-site
USD 90,000 - 100,000
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Equity
Benefits
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA • Austin (TX)

On-site
USD 124,000 - 196,000
Equity
Benefits
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA • Westford (MA)

On-site
USD 124,000 - 196,000
Equity