HPC Engineer

Ifm Us

Sunnyvale (CA)

On-site

USD 80,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
Generous paid time off
Paid Parental Leave
Employee Assistance Program
Life insurance and disability

Job summary

Ifm Us is seeking an individual for a role that provides operational coverage during overnight hours, primarily for monitoring large-scale GPU clusters and supporting researchers. Responsibilities include incident response, troubleshooting job failures, and validating deployments.

Candidates should have a Bachelor’s degree and 2+ years of experience in relevant fields including Linux systems administration and scripting. Benefits include comprehensive medical, dental and vision coverage, paid time off, and more.

Qualifications

  • 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.
  • Strong Linux troubleshooting skills.
  • Experience with scripting using Python or Bash.

Responsibilities

  • Monitor health, performance, and availability of large-scale GPU clusters.
  • Respond to incidents and perform first-level triage.
  • Support researchers and troubleshoot job failures.
  • Execute operational runbooks and recovery procedures.
  • Validate cluster deployments, upgrades, and maintenance activities.
  • Track infrastructure utilization and operational metrics.
  • Develop automation and monitoring tools.
  • Contribute to documentation and reporting.

Skills

Linux troubleshooting
Scripting (Python or Bash)
Slurm
GPU infrastructure
AWS
Azure
GCP
Grafana
Prometheus
Datadog
Containers
Kubernetes
AI/ML infrastructure
Research computing environments

Education

Bachelor's degree in relevant field

Job description

About MBZUAI

The Institute for Foundation Models (IFM) operates some of the world's largest AI supercomputing environments.

Position Summary

This role provides operational coverage during Abu Dhabi overnight hours and serves as a primary point of contact for infrastructure monitoring, incident triage, researcher support, and production operations.

Responsibilities
  • Monitor health, performance, and availability of large-scale GPU clusters.
  • Respond to incidents and perform first-level triage.
  • Support researchers and troubleshoot job failures.
  • Execute operational runbooks and recovery procedures.
  • Validate cluster deployments, upgrades, and maintenance activities.
  • Track infrastructure utilization and operational metrics.
  • Develop automation and monitoring tools.
  • Contribute to documentation and reporting.
Qualifications

Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, Information Technology, Electrical Engineering, Mathematics, Physics, or related disciplines.

  • 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.
  • Strong Linux troubleshooting skills.
  • Experience with scripting using Python or Bash.
  • Slurm.
  • GPU infrastructure.
  • AWS, Azure, or GCP.
  • Grafana, Prometheus, Datadog, or similar tools.
  • Containers and Kubernetes.
  • AI/ML infrastructure exposure.
  • Research computing environments.
Benefits Include
  • Comprehensive medical, dental, and vision benefits
  • Bonus
  • 401K Plan
  • Generous paid time off, sick leave and holidays
  • Paid Parental Leave
  • Employee Assistance Program
  • Life insurance and disability
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Engineer
HPC Engineer

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 300,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
Overnight HPC Operations Engineer – GPU, Slurm
Overnight HPC Operations Engineer – GPU, Slurm

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 300,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
Overnight HPC Engineer — AI Research Infra
Overnight HPC Engineer — AI Research Infra

Ifm Us • Sunnyvale (CA)

On-site
USD 80,000 - 110,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Research Scientist - Distributed Machine Learning
Research Scientist - Distributed Machine Learning

Ifm Us • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Health benefits
401K Plan
Paid time off
+2
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer
AI and ML HPC Cluster Engineer, AI and ML HPC Cluster Engineer

NVIDIA • Colorado

On-site
USD 124,000 - 196,000
AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

Remote
USD 110,000 - 150,000
Professional development
Conferences attendance
Team events
HPC Administrator
HPC Administrator

Jahnel Group • United States

On-site
USD 120,000 - 180,000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • California (MO)

On-site
USD 176,000 - 334,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000