AI Infra Engineer: GPU Fleet Automation

Fal.ai Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 250,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Relocation assistance
Health, dental, and vision insurance (
Regular team events and offsites

Job summary

fal.ai Inc. seeks a hands-on engineer to keep a large fleet of GPU servers healthy and productive in San Francisco, CA.

You will design and operate software for provisioning, health monitoring, error detection, and automated recovery across thousands of servers, with emphasis on reliability and scalable Linux systems. This role offers $180,000-$250,000 per year plus equity and benefits, and relocation assistance to San Francisco.

Qualifications

  • 3+ years experience managing bare-metal and cloud based server fleets at scale (100+ nodes)
  • Strong software engineering skills in Python; you write production tooling, not scripts
  • Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling
  • Strong experience with configuration management and infrastructure-as-code: Ansible, Terraform, cloud-init
  • Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning
  • Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory)
  • Experience building internal tools or dashboards for infrastructure visibility
  • Excellent communication and ability to drive technical decisions across teams
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement

Responsibilities

  • Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc
  • Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting
  • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals)
  • Leverage AI to an extreme level to build tools and automate alerting and recovery
  • Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation
  • Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and temporary scratch: NVMe arrays, NFS, parallel file systems, and object storage
  • Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes)
  • Develop a suite of automated error detection and recovery processes
  • Work with partners to solve technical issues

Skills

Python
Bare-metal fleets
Linux
Ansible
Terraform
cloud-init
Storage technologies
NVMe
GPU diagnostics
Security hardening
CI/CD dashboards

Tools

PXE/iPXE
Kickstart
libvirt
Qemu/KVM
NVIDIA DCGM

Job description

fal.ai Inc. seeks a hands-on engineer to keep a large fleet of GPU servers healthy and productive in San Francisco, CA.

You will design and operate software for provisioning, health monitoring, error detection, and automated recovery across thousands of servers, with emphasis on reliability and scalable Linux systems. This role offers $180,000-$250,000 per year plus equity and benefits, and relocation assistance to San Francisco.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Infra Engineer: Python tooling for AI server fleets
GPU Infra Engineer: Python tooling for AI server fleets

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Health, dental, and vision insurance (
Team events and offsites
+1
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery

Together • San Francisco (CA)

On-site
USD 190,000 - 270,000
Health insurance
Startup equity
Competitive benefits
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Generative AI Infra Engineer (GPU Fleet)
Generative AI Infra Engineer (GPU Fleet)

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
Staff Engineer, GPU AI Inference & RL Infrastructure
Staff Engineer, GPU AI Inference & RL Infrastructure

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Staff AI Infra Engineer: GPU Fleet Reliability Leader
Staff AI Infra Engineer: GPU Fleet Reliability Leader

Luma AI • United States

Remote
USD 210,000 - 320,000
GPU Infra Engineer — Automate & Scale Hyperscale AI Compute
GPU Infra Engineer — Automate & Scale Hyperscale AI Compute

Fluidstack • San Francisco (CA)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Generous PTO policy
Retirement or pension plan
GPU & Compute Infra Engineer — Remote/Hybrid
GPU & Compute Infra Engineer — Remote/Hybrid

Lightning AI • San Francisco (CA), Seattle (WA), New York (NY)

Hybrid
USD 180,000 - 200,000
Comprehensive medical, dental, and vision coverage
Generous paid time off
Flexible work environment
Software Engineer, Platform
Software Engineer, Platform

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Health, dental, and vision insurance (
Team events and offsites
+1