Overnight HPC Operations Engineer – GPU, Slurm

Institute of Foundation Models

Sunnyvale (CA)

On-site

USD 150,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
Generous paid time off
Paid Parental Leave
Employee Assistance Program
Life insurance and disability

Job summary

The Institute of Foundation Models in Sunnyvale, CA is hiring for a role focused on operational coverage during Abu Dhabi overnight hours. This position involves monitoring the health and performance of large-scale GPU clusters, responding to incidents, and supporting researchers with troubleshooting.

The ideal candidate will require a Bachelor's degree in a relevant field and at least 2 years of experience in Linux systems administration or related fields. Attractive benefits include comprehensive health coverage, bonuses, and a 401K Plan.

Qualifications

  • 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.
  • Experience with scripting using Python or Bash.
  • Exposure to AI/ML infrastructure and research computing environments.

Responsibilities

  • Monitor health, performance, and availability of large-scale GPU clusters.
  • Respond to incidents and perform first-level triage.
  • Support researchers and troubleshoot job failures.
  • Execute operational runbooks and recovery procedures.
  • Validate cluster deployments, upgrades, and maintenance activities.
  • Contribute to documentation and reporting.

Skills

Linux systems administration
Scripting (Python or Bash)
Strong Linux troubleshooting skills

Education

Bachelor's degree in a relevant field

Tools

Slurm
GPU infrastructure
AWS
Azure
GCP
Grafana
Prometheus
Datadog
Kubernetes

Job description

The Institute of Foundation Models in Sunnyvale, CA is hiring for a role focused on operational coverage during Abu Dhabi overnight hours. This position involves monitoring the health and performance of large-scale GPU clusters, responding to incidents, and supporting researchers with troubleshooting.

The ideal candidate will require a Bachelor's degree in a relevant field and at least 2 years of experience in Linux systems administration or related fields. Attractive benefits include comprehensive health coverage, bonuses, and a 401K Plan.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Overnight HPC Engineer — AI Research Infra
Overnight HPC Engineer — AI Research Infra

Ifm Us • Sunnyvale (CA)

On-site
USD 80,000 - 110,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
HPC Engineer
HPC Engineer

Ifm Us • Sunnyvale (CA)

On-site
USD 80,000 - 110,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
HPC Engineer
HPC Engineer

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 300,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
Staff HPC Systems Engineer: Slurm Architect for GPU Cloud
Staff HPC Systems Engineer: Slurm Architect for GPU Cloud

Nscale • New York (NY)

On-site
USD 225,000 - 275,000
Highly competitive compensation package
Performance reviews every 12 months
Flexible workplace
Senior HPC GPU Compute Engineer (Hybrid SF)
Senior HPC GPU Compute Engineer (Hybrid SF)

The San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Generous equity grant
Retirement matching
Comprehensive medical, dental, and vision insurance
+3
Senior HPC Support Engineer - GPU Cloud Infra
Senior HPC Support Engineer - GPU Cloud Infra

Neura Market • United States

On-site
USD 150,000 - 190,000
Wellness stipend
Commuter stipend
401k plan with 2% company match (USA)
+1
Senior GPU HPC Cluster Engineer — Equity Eligible
Senior GPU HPC Cluster Engineer — Equity Eligible

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Lead HPC Cluster Engineer for GPU AI Compute
Lead HPC Cluster Engineer for GPU AI Compute

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6