GPU Cluster Infra Engineer - Reliability & Automation

Doist

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Health benefits
401k matching
Unlimited PTO
Refill Days

Job summary

Liquid AI is seeking a hands-on software engineer to own the reliability and operation of GPU clusters used for training and research. You will debug issues across compute, storage, networking, schedulers, and distributed workloads, while improving CPU/GPU storage utilization through tooling and automation.

You will onboard and migrate workloads across providers and hardware platforms, build monitoring and platform abstractions, and contribute to the longer-term architecture of our training

Qualifications

  • Strong production-grade infra tooling and automation
  • Deep knowledge of distributed systems, Linux, networking, storage
  • Experience operating a shared compute cluster or distributed training platform
  • Proven ability to turn recurring issues into durable solutions
  • Ability to partner with senior research and infrastructure engineers

Responsibilities

  • Own the reliability and operation of GPU clusters for training/research
  • Debug issues across compute, storage, networking, schedulers, and distributed workloads
  • Improve utilization via better tooling and automation
  • Onboard and migrate workloads across GPU providers and hardware platforms
  • Build monitoring, validation, and platform abstractions to reduce researcher workload
  • Contribute to longer-term architecture of training infrastructure and GPU platform

Skills

Software engineering
Distributed systems
Linux
Networking
Storage
Automation
Monitoring

Tools

SLURM
Kubernetes
Ray
Hadoop

Job description

Liquid AI is seeking a hands-on software engineer to own the reliability and operation of GPU clusters used for training and research. You will debug issues across compute, storage, networking, schedulers, and distributed workloads, while improving CPU/GPU storage utilization through tooling and automation.

You will onboard and migrate workloads across providers and hardware platforms, build monitoring and platform abstractions, and contribute to the longer-term architecture of our training

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
GPU Cluster Architect: Scalable AI Platform
GPU Cluster Architect: Scalable AI Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
Senior GPU Infra Engineer for Distributed AI
Senior GPU Infra Engineer for Distributed AI

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - GPU Infrastructure Engineer
Member of Technical Staff - GPU Infrastructure Engineer

Liquid AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
Health benefits
401k matching
+2
Staff AI Infra Engineer: GPU Fleet Reliability Leader
Staff AI Infra Engineer: GPU Fleet Reliability Leader

Luma AI • United States

Remote
USD 210,000 - 320,000
GPU Infrastructure Engineer - Scale & Automation
GPU Infrastructure Engineer - Scale & Automation

United States Digital Space LLC • United States

Remote
USD 150,000 - 210,000
Senior AI Training Infra Engineer - Scale GPU Clusters
Senior AI Training Infra Engineer - Scale GPU Clusters

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
401(k) with company match
Paid holidays
GPU Cluster Engineer - Scalable AI Infrastructure
GPU Cluster Engineer - Scalable AI Infrastructure

NEURA Robotics • Germany (OH)

On-site
USD 140,000 - 195,000
Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2