Remote AI Research Clusters Engineer - ML Infra & GPU

NEPSE Trading

Northern (KY)

Hybrid

USD 124,000 - 196,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NVIDIA is seeking a Software Engineer to advance AI research clusters on scalable GPU systems. You will design and implement solutions to deploy reliable, secure, and high-performance GPU clusters for internal researchers, enabling them to train, fine-tune, and deploy ML models.

You will work with the AI Platform to reduce operational toil, perform on-call support, and explore AIOps and Agentic AI in production. The role requires strong Linux, Python, C++, and distributed systems skills.

Qualifications

  • BS/MS in Computer Science, Engineering, or equivalent experience.
  • 2+ years in software/platform engineering, including 1 year in ML infrastructure or distributed systems.
  • Experience in software development lifecycle on Linux-based platforms.
  • Strong coding skills in Python, C++ or Rust.
  • Experience with Docker, Kubernetes, GitLab CI, automated deployments.
  • Experience with AIOps or Agentic AI and apply it successfully in production environment.

Responsibilities

  • Understand pain points of validating, monitoring and operating GPU clusters at scale.
  • Design, develop and maintain engineering solutions to solve those pain points.
  • Research in traditional AIOps and Agentic AI to reduce operation toil.
  • Participate in on-call support for systems, platforms built and owned by the team.

Job description

NVIDIA is seeking a Software Engineer to advance AI research clusters on scalable GPU systems. You will design and implement solutions to deploy reliable, secure, and high-performance GPU clusters for internal researchers, enabling them to train, fine-tune, and deploy ML models.

You will work with the AI Platform to reduce operational toil, perform on-call support, and explore AIOps and Agentic AI in production. The role requires strong Linux, Python, C++, and distributed systems skills.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Research Clusters Engineer — GPU ML Infra & AIOps
AI Research Clusters Engineer — GPU ML Infra & AIOps

NVIDIA AI • Durham (NC)

On-site
USD 120,000 - 180,000
Equity
Benefits
Senior AI Performance & Efficiency Engineer - Equity Eligible
Senior AI Performance & Efficiency Engineer - Equity Eligible

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Competitive benefits
AI/ML Infra Engineer — GPU Clusters & HPC
AI/ML Infra Engineer — GPU Clusters & HPC

NVIDIA AI • Redmond (WA)

On-site
USD 120,000 - 180,000
Competitive salaries
Comprehensive benefits package
Equity
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA AI • Durham (NC)

On-site
USD 120,000 - 180,000
Equity
Benefits
Remote ML Infrastructure Engineer — GPU Clusters & AI Platform
Remote ML Infrastructure Engineer — GPU Clusters & AI Platform

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior AI Infra Engineer-Distributed GPU Clusters (Equity)
Senior AI Infra Engineer-Distributed GPU Clusters (Equity)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Remote AI Systems Engineer - Scale GPU ML Infra
Remote AI Systems Engineer - Scale GPU ML Infra

Bright-Vision-Technologies • United States

Remote
USD 90,000 - 100,000
ML Systems Engineer - Distributed AI & GPU
ML Systems Engineer - Distributed AI & GPU

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits package