GPU Cluster Infrastructure Engineer

PVH (Tommy Hilfiger/Calvin Klein)

United States

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity in unicorn-stage company
100% premiums covered for medical, den
401(k) matching up to 4%
Unlimited PTO
Refill Days

Job summary

Liquid AI is seeking a hands-on software engineer to own the reliability of GPU clusters used for foundation model training and research, and to build tooling that increases resource utilization and reduces toil for researchers.

You will work closely with researchers and infrastructure engineers, addressing issues from immediate operational response to long‑term platform improvements, with a strong emphasis on automation and monitoring.

Qualifications

  • Strong software engineering experience with production-quality infrastructure tooling and automation.
  • Deep knowledge of distributed systems, Linux, networking, and storage.
  • Experience operating a shared compute cluster or distributed training platform.
  • Track record of supporting production users and turning recurring failures into durable solutions.

Responsibilities

  • Own the reliability and operation of the GPU clusters used for training and research.
  • Debug issues across compute, storage, networking, schedulers, and distributed workloads.
  • Improve CPU, GPU, and storage utilization through better tooling and automation.
  • Onboard and migrate workloads across GPU providers and hardware platforms.
  • Build monitoring, validation, and platform abstractions that reduce operational work for researchers.
  • Contribute to the longer‑term architecture of Liquid AI's training infrastructure and GPU platform.

Skills

Software engineering
Distributed systems
Linux
Networking
Storage
Automation
Production users support
Problem solving

Tools

SLURM
Kubernetes
Ray
Hadoop
Distributed storage
Cloud providers

Job description

Liquid AI is seeking a hands-on software engineer to own the reliability of GPU clusters used for foundation model training and research, and to build tooling that increases resource utilization and reduces toil for researchers.

You will work closely with researchers and infrastructure engineers, addressing issues from immediate operational response to long‑term platform improvements, with a strong emphasis on automation and monitoring.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Cluster Infra Engineer - Reliability & Automation
GPU Cluster Infra Engineer - Reliability & Automation

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
Health benefits
401k matching
+2
GPU Infrastructure Engineer — Scale AI Clusters, Equity
GPU Infrastructure Engineer — Scale AI Clusters, Equity

Liquid AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health insurance
401(k) matching up to 4%
Unlimited PTO
+2
Member of Technical Staff - GPU Infrastructure Engineer
Member of Technical Staff - GPU Infrastructure Engineer

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
Health benefits
401k matching
+2
Member of Technical Staff - GPU Infrastructure Engineer
Member of Technical Staff - GPU Infrastructure Engineer

Liquid AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health insurance
401(k) matching up to 4%
Unlimited PTO
+2
Member of Technical Staff - GPU Infrastructure Engineer
Member of Technical Staff - GPU Infrastructure Engineer

Liquid AI • United States

On-site
USD 140,000 - 210,000
Equity in unicorn-stage company
100% premiums covered for medical, den
401(k) matching up to 4%
+2
GPU Cluster Architect: Scalable AI Platform
GPU Cluster Architect: Scalable AI Platform

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 270,000
Medical insurance
401k plan
Daily meals/snacks
+2
Lead AI Infrastructure Engineer: GPU Clusters & Reliability
Lead AI Infrastructure Engineer: GPU Clusters & Reliability

Luma AI • San Francisco (CA)

On-site
USD 300,000 - 420,000
Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Staff Engineer, Distributed GPU Clusters
Staff Engineer, Distributed GPU Clusters

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Staff GPU Compute Infrastructure Engineer
Staff GPU Compute Infrastructure Engineer

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000