GPU ML Infra Engineer for Clusters & AIOps Equity

Socket.dev

North Carolina

On-site

USD 124,000 - 196,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Software Engineer to accelerate the next era of machine learning innovation, focusing on delivering reliable GPU clusters to internal researchers and minimizing operational disruption.

You will design and implement engineering solutions, explore AIOps and Agentic AI to reduce toil, and participate in on-call support for systems built by the team. Strong Linux, Python/C++/Rust skills are expected.

Qualifications

  • BS/MS in Computer Science, Engineering, or equivalent experience.
  • 2+ years in software/platform engineering, including 1 year in ML infrastructure or distributed systems.
  • Experience in software development lifecycle on Linux-based platforms.
  • Strong coding skills in languages such as Python, C++ or Rust.
  • Experience with Docker, Kubernetes, GitLab CI, automated deployments.
  • Experience with AIOps or Agentic AI and apply it successfully in production environment.

Responsibilities

  • Work with the AI Platform org to understand pain points of validating, monitoring and operating GPU clusters at scale, then design and maintain engineering solutions.
  • Research in traditional AIOps and Agentic AI to reduce operation toil.
  • Participate in on-call support for systems and platforms owned by the team.

Skills

Python
C++
Rust
AIOps awareness

Education

BS/MS in Computer Science or Engineering

Tools

Docker
Kubernetes
GitLab CI
CI/CD

Job description

NVIDIA is seeking a Software Engineer to accelerate the next era of machine learning innovation, focusing on delivering reliable GPU clusters to internal researchers and minimizing operational disruption.

You will design and implement engineering solutions, explore AIOps and Agentic AI to reduce toil, and participate in on-call support for systems built by the team. Strong Linux, Python/C++/Rust skills are expected.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infra Engineer - GPU Clusters & AIOps
ML Infra Engineer - GPU Clusters & AIOps

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 124,000 - 196,000
ML Infra Engineer: GPU Clusters & Self-Serve Scale
ML Infra Engineer: GPU Clusters & Self-Serve Scale

NVIDIA • Austin (TX)

On-site
USD 124,000 - 196,000
Equity
Benefits
GPU AI Research Clusters Engineer — Scale ML Infra
GPU AI Research Clusters Engineer — Scale ML Infra

NVIDIA • Westford (MA)

On-site
USD 124,000 - 196,000
Equity
AI Research Clusters Engineer — GPU ML Infra & AIOps
AI Research Clusters Engineer — GPU ML Infra & AIOps

NVIDIA AI • Durham (NC)

On-site
USD 120,000 - 180,000
Equity
Benefits
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA • Westford (MA)

On-site
USD 124,000 - 196,000
Equity
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA • Austin (TX)

On-site
USD 124,000 - 196,000
Equity
Benefits
Senior AI Performance & Efficiency Engineer - Equity Eligible
Senior AI Performance & Efficiency Engineer - Equity Eligible

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Equity
Benefits
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 124,000 - 196,000
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA AI • Durham (NC)

On-site
USD 120,000 - 180,000
Equity
Benefits