HPC Engineer

Institute of Foundation Models

Sunnyvale (CA)

On-site

USD 150,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
Generous paid time off
Paid Parental Leave
Employee Assistance Program
Life insurance and disability

Job summary

The Institute of Foundation Models in Sunnyvale, CA is hiring for a role focused on operational coverage during Abu Dhabi overnight hours. This position involves monitoring the health and performance of large-scale GPU clusters, responding to incidents, and supporting researchers with troubleshooting.

The ideal candidate will require a Bachelor's degree in a relevant field and at least 2 years of experience in Linux systems administration or related fields. Attractive benefits include comprehensive health coverage, bonuses, and a 401K Plan.

Qualifications

  • 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.
  • Experience with scripting using Python or Bash.
  • Exposure to AI/ML infrastructure and research computing environments.

Responsibilities

  • Monitor health, performance, and availability of large-scale GPU clusters.
  • Respond to incidents and perform first-level triage.
  • Support researchers and troubleshoot job failures.
  • Execute operational runbooks and recovery procedures.
  • Validate cluster deployments, upgrades, and maintenance activities.
  • Contribute to documentation and reporting.

Skills

Linux systems administration
Scripting (Python or Bash)
Strong Linux troubleshooting skills

Education

Bachelor's degree in a relevant field

Tools

Slurm
GPU infrastructure
AWS
Azure
GCP
Grafana
Prometheus
Datadog
Kubernetes

Job description

About MBZUAI

The Institute for Foundation Models (IFM) operates some of the world's largest AI supercomputing environments.

Position Summary

This role provides operational coverage during Abu Dhabi overnight hours and serves as a primary point of contact for infrastructure monitoring, incident triage, researcher support, and production operations.

Responsibilities
  • Monitor health, performance, and availability of large-scale GPU clusters.
  • Respond to incidents and perform first-level triage.
  • Support researchers and troubleshoot job failures.
  • Execute operational runbooks and recovery procedures.
  • Validate cluster deployments, upgrades, and maintenance activities.
  • Track infrastructure utilization and operational metrics.
  • Develop automation and monitoring tools.
  • Contribute to documentation and reporting.
Education

Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, Information Technology, Electrical Engineering, Mathematics, Physics, or related disciplines.

Experience
  • 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.
  • Strong Linux troubleshooting skills.
  • Experience with scripting using Python or Bash.
Preferred Qualifications
  • Slurm.
  • GPU infrastructure.
  • AWS, Azure, or GCP.
  • Grafana, Prometheus, Datadog, or similar tools.
  • Containers and Kubernetes.
  • AI/ML infrastructure exposure.
  • Research computing environments.
Salary Range

$150,000 - $300,000 a year

The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs.

Benefits Include
  • Comprehensive medical, dental, and vision benefits
  • Bonus
  • 401K Plan
  • Generous paid time off, sick leave and holidays
  • Paid Parental Leave
  • Employee Assistance Program
  • Life insurance and disability
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Engineer
HPC Engineer

Ifm Us • Sunnyvale (CA)

On-site
USD 80,000 - 110,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
Overnight HPC Operations Engineer – GPU, Slurm
Overnight HPC Operations Engineer – GPU, Slurm

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 300,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
Distributed Machine Learning Engineer
Distributed Machine Learning Engineer

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
Overnight HPC Engineer — AI Research Infra
Overnight HPC Engineer — AI Research Infra

Ifm Us • Sunnyvale (CA)

On-site
USD 80,000 - 110,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer

NVIDIA • Colorado

On-site
USD 124,000 - 196,000
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

On-site
USD 176,000 - 334,000
Research Scientist - Distributed Machine Learning
Research Scientist - Distributed Machine Learning

Ifm Us • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Health benefits
401K Plan
Paid time off
+2
IT Operations Specialist
IT Operations Specialist

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 100,000 - 150,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4