AI Infrastructure Engineer: GPU Fleet & Automation

fal

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Relocation assistance to San Francisco
Health, dental, and vision insurance (
Team events and offsites

Job summary

fal is seeking a hands-on engineer to build and operate the software and processes that keep a large fleet of GPU servers healthy and productive. You will write systems and tooling for provisioning, health monitoring, error detection, and recovery, and drive resolution with partners when automation fails.

You will own Python fleet tooling, implement OS security baselines, optimize storage, and tune Linux for AI workloads.

Qualifications

  • 3+ years managing bare-metal and cloud server fleets at scale.
  • Strong Python production tooling experience.
  • Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, and I/O stack.
  • Experience with Ansible, Terraform, cloud-init.
  • Storage technologies: LVM, RAID, NVMe, NFS, Lustre/GPFS.
  • Familiarity with hardware diagnostics for GPUs/NVMe/NICs/memory.
  • Experience building internal tools or dashboards for infra visibility.
  • Excellent communication and ability to drive tech decisions across teams.
  • Self-starter who executes quickly and takes ownership.

Responsibilities

  • Build and maintain Python fleet tracking system that manages full lifecycle of servers.
  • Build server management tooling for provisioning, health checks, GPU diagnostics, recovery and alerting.
  • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet.
  • Leverage AI to build tools and automate alerting and recovery.
  • Implement and enforce OS-level security baselines and policies.
  • Manage and optimize storage systems supporting model weights and scratch data.
  • Tune Linux systems for AI workloads (kernel, NUMA, CPU pinning, I/O schedulers).
  • Develop automated error detection and recovery processes.
  • Work with partners to solve technical issues.

Skills

Python
Linux
SRE tooling
Terraform
Ansible
Networking
Diagnostics
Communication
Ownership

Tools

NVMe
NFS
GPFS

Job description

fal is seeking a hands-on engineer to build and operate the software and processes that keep a large fleet of GPU servers healthy and productive. You will write systems and tooling for provisioning, health monitoring, error detection, and recovery, and drive resolution with partners when automation fails.

You will own Python fleet tooling, implement OS security baselines, optimize storage, and tune Linux for AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Generative AI Infra Engineer (GPU Fleet)
Generative AI Infra Engineer (GPU Fleet)

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
GPU Infra Engineer: Python tooling for AI server fleets
GPU Infra Engineer: Python tooling for AI server fleets

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Health, dental, and vision insurance (
Team events and offsites
+1
AI Infra Engineer: GPU Fleet Automation
AI Infra Engineer: GPU Fleet Automation

Fal.ai Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Regular team events and offsites
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery

Together • San Francisco (CA)

On-site
USD 190,000 - 270,000
Health insurance
Startup equity
Competitive benefits
AI Infrastructure & Automation Engineer
AI Infrastructure & Automation Engineer

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities
GPU Fleet Orchestrator for AI Infrastructure
GPU Fleet Orchestrator for AI Infrastructure

AMD • San Jose (CA)

Hybrid
USD 180,000 - 240,000
Staff AI Infra Engineer: GPU Fleet Reliability Leader
Staff AI Infra Engineer: GPU Fleet Reliability Leader

Luma AI • United States

Remote
USD 210,000 - 320,000
GPU Compute Engineer — Fleet Reliability & Automation
GPU Compute Engineer — Fleet Reliability & Automation

Fluidstack • San Francisco (CA)

On-site
USD 175,000 - 300,000
Health, dental, and vision insurance
Retirement or pension plan
Generous PTO policy
AI Compute Fleet Engineer – GPU & Server Automation
AI Compute Fleet Engineer – GPU & Server Automation

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 149,000 - 201,000
Health insurance
401(k) matching
Paid time off
GPU Fleet Orchestrator for AI Infra
GPU Fleet Orchestrator for AI Infra

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 150,000 - 210,000