AI Research Clusters Engineer — GPU ML Infra & AIOps

NVIDIA AI

Durham (NC)

On-site

USD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA AI in Durham, NC is seeking an engineer to design and implement solutions that optimize the reliability and performance of GPU clusters for internal AI researchers. You will tackle complex ML infrastructure challenges and improve tooling for researchers.

Responsibilities include reducing operational toil through AIOps and Agentic AI, and providing on-call support for platforms. Requires a BS/MS in CS or Engineering with 2+ years in software engineering and proficiency in Python, C++, or

Qualifications

  • Requires a BS/MS in CS or engineering with 2+ years in software engineering, including ML infrastructure.
  • Proficiency in Python, C++, or Rust and experience with Docker and Kubernetes.

Responsibilities

  • Design and implement engineering solutions to optimize GPU cluster reliability and performance.
  • Provide on-call support for platforms and reduce operational toil with AIOps.

Skills

Python
C++
Rust
Docker
Kubernetes
Linux
Distributed Systems
ML Infrastructure
REST API
Slurm
GitLab CI

Education

Bachelor's degree in Computer Science or Engineering
Master's degree

Tools

Docker
Kubernetes
Slurm
GitLab CI

Job description

NVIDIA AI in Durham, NC is seeking an engineer to design and implement solutions that optimize the reliability and performance of GPU clusters for internal AI researchers. You will tackle complex ML infrastructure challenges and improve tooling for researchers.

Responsibilities include reducing operational toil through AIOps and Agentic AI, and providing on-call support for platforms. Requires a BS/MS in CS or Engineering with 2+ years in software engineering and proficiency in Python, C++, or

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU ML Infra Engineer for Clusters & AIOps Equity
GPU ML Infra Engineer for Clusters & AIOps Equity

Socket.dev • North Carolina

On-site
USD 124,000 - 196,000
Equity
Benefits
ML Infra Engineer - GPU Clusters & AIOps
ML Infra Engineer - GPU Clusters & AIOps

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 124,000 - 196,000
ML Infra Engineer: GPU Clusters & Self-Serve Scale
ML Infra Engineer: GPU Clusters & Self-Serve Scale

NVIDIA • Austin (TX)

On-site
USD 124,000 - 196,000
Equity
Benefits
AI/ML Infra Engineer — GPU Clusters & HPC
AI/ML Infra Engineer — GPU Clusters & HPC

NVIDIA AI • Redmond (WA)

On-site
USD 120,000 - 180,000
Competitive salaries
Comprehensive benefits package
Equity
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA AI • Durham (NC)

On-site
USD 120,000 - 180,000
Equity
Benefits
Senior AI Performance & Efficiency Engineer - Equity Eligible
Senior AI Performance & Efficiency Engineer - Equity Eligible

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
Lead AI Infrastructure Architect for Large-Scale GPU Clusters
Lead AI Infrastructure Architect for Large-Scale GPU Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA • Austin (TX)

On-site
USD 124,000 - 196,000
Equity
Benefits
AI Data Center Field Engineer: GPU & Networking
AI Data Center Field Engineer: GPU & Networking

NVIDIA • Raleigh (NC)

On-site
USD 132,000 - 207,000
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package