Software Engineer - AI Research Clusters

NVIDIA AI

Durham (NC)

On-site

USD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA AI in Durham, NC is seeking an engineer to design and implement solutions that optimize the reliability and performance of GPU clusters for internal AI researchers. You will tackle complex ML infrastructure challenges and improve tooling for researchers.

Responsibilities include reducing operational toil through AIOps and Agentic AI, and providing on-call support for platforms. Requires a BS/MS in CS or Engineering with 2+ years in software engineering and proficiency in Python, C++, or

Qualifications

  • Requires a BS/MS in CS or engineering with 2+ years in software engineering, including ML infrastructure.
  • Proficiency in Python, C++, or Rust and experience with Docker and Kubernetes.

Responsibilities

  • Design and implement engineering solutions to optimize GPU cluster reliability and performance.
  • Provide on-call support for platforms and reduce operational toil with AIOps.

Skills

Python
C++
Rust
Docker
Kubernetes
Linux
Distributed Systems
ML Infrastructure
REST API
Slurm
GitLab CI

Education

Bachelor's degree in Computer Science or Engineering
Master's degree

Tools

Docker
Kubernetes
Slurm
GitLab CI

Job description

Design and implement engineering solutions to optimize the reliability and performance of GPU clusters for internal AI researchers. This includes reducing operational toil through AIOps and Agentic AI while providing on-call support for the platforms.

Requirements: Requires a BS/MS in Computer Science or Engineering with 2+ years of software engineering experience, including at least one year in ML infrastructure. Proficiency in Python, C++, or Rust and experience with containerization tools like Docker and Kubernetes is essential.

Key Skills: Python, C++, Rust, Docker, Kubernetes, GitLab CI, AIOps, Agentic AI, Linux, Distributed Systems, ML Infrastructure, Full-stack Development, Relational Data Modeling, REST API, Slurm, GPU Computing

Benefits: Equity, Benefits

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA • Austin (TX)

On-site
USD 124,000 - 196,000
Equity
Benefits
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 124,000 - 196,000
GPU ML Infra Engineer for Clusters & AIOps Equity
GPU ML Infra Engineer for Clusters & AIOps Equity

Socket.dev • North Carolina

On-site
USD 124,000 - 196,000
Equity
Benefits
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

Socket.dev • North Carolina

On-site
USD 124,000 - 196,000
Equity
Benefits
AI Research Clusters Engineer — GPU ML Infra & AIOps
AI Research Clusters Engineer — GPU ML Infra & AIOps

NVIDIA AI • Durham (NC)

On-site
USD 120,000 - 180,000
Equity
Benefits
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
AI/ML Infra Engineer — GPU Clusters & HPC
AI/ML Infra Engineer — GPU Clusters & HPC

NVIDIA AI • Redmond (WA)

On-site
USD 120,000 - 180,000
Competitive salaries
Comprehensive benefits package
Equity
AI and ML Infra Software Engineer, GPU Clusters - New College Grad 2026
AI and ML Infra Software Engineer, GPU Clusters - New College Grad 2026

NVIDIA AI • Redmond (WA)

On-site
USD 120,000 - 180,000
Competitive salaries
Comprehensive benefits package
Equity
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

Cursor • California (MO)

On-site
USD 140,000 - 185,000