Senior GPU Fleet Reliability Engineer

Fal

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision insurance
Learning and growth opportunities
Visa sponsorship and relocation assistance
Regular team events and offsites

Job summary

A tech company located in San Francisco seeks a hands-on engineer to manage a fleet of GPU servers. The role requires building and maintaining a Python fleet tracking system, implementing OS-level security, and developing automated error detection processes. The ideal candidate should have over 5 years of experience managing server fleets, with strong Python programming skills and a deep understanding of Linux systems. Compensation includes a salary range of $180,000-250,000, plus equity and benefits.

Qualifications

  • 5+ years experience managing bare-metal and VM server fleets at scale.
  • Strong software engineering skills in Python; writing production tooling.
  • Deep understanding of Linux systems including boot process and networking.

Responsibilities

  • Build and maintain Python fleet tracking system for server lifecycle management.
  • Implement OS-level security and manage distributed storage systems.
  • Develop automated error detection and recovery processes.

Skills

Python
Linux systems knowledge
Configuration management
Infrastructure-as-code
Storage technologies
Communication skills
Self-starter

Tools

Ansible
Terraform

Job description

A tech company located in San Francisco seeks a hands-on engineer to manage a fleet of GPU servers. The role requires building and maintaining a Python fleet tracking system, implementing OS-level security, and developing automated error detection processes. The ideal candidate should have over 5 years of experience managing server fleets, with strong Python programming skills and a deep understanding of Linux systems. Compensation includes a salary range of $180,000-250,000, plus equity and benefits.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Engineer, GPU Fleet Automation & Ops
Senior Engineer, GPU Fleet Automation & Ops

NVIDIA • New York (NY)

On-site
USD 308,000 - 472,000
Senior GPU HPC Platform Reliability Engineer
Senior GPU HPC Platform Reliability Engineer

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 590,000
GPU Reliability Engineer - AI Supercomputing Fleet
GPU Reliability Engineer - AI Supercomputing Fleet

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Unlimited PTO
Paid parental leave
Relocation support
+1
Software Engineer - GPU Fleet
Software Engineer - GPU Fleet

Iceberg • New York (NY)

On-site
USD 120,000 - 170,000
GPU HPC Fleet Reliability Engineer
GPU HPC Fleet Reliability Engineer

CoreWeave • Bellevue (WA)

Hybrid
USD 83,000 - 110,000
Software Engineer, Infrastructure
Software Engineer, Infrastructure

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Health, dental, and vision insurance
Learning and growth opportunities
Visa sponsorship and relocation assistance
+1
Senior HPC GPU Compute Engineer (Hybrid SF)
Senior HPC GPU Compute Engineer (Hybrid SF)

The San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Generous equity grant
Retirement matching
Comprehensive medical, dental, and vision insurance
+3
Head of GPU Fleet Automation & Infrastructure Engineering
Head of GPU Fleet Automation & Infrastructure Engineering

nscaleoperationsukltd • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior GPU HPC Systems Engineer
Senior GPU HPC Systems Engineer

Harvey Nash • Chicago (IL)

Hybrid
USD 125,000 - 150,000
Fleet Reliability Engineer — HPC & GPU Clusters
Fleet Reliability Engineer — HPC & GPU Clusters

CoreWeave • Plano (TX)

Hybrid
USD 83,000 - 110,000