GPU Infra Engineer for Scalable AI Platform

Kindredventures

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Relocation to San Francisco
Health, dental, and vision insurance (
Regular team events
Learning and growth opportunities

Job summary

fal is building a scalable generative AI infrastructure. We seek an hands-on engineer to manage thousands of servers, automate provisioning, health checks, and recovery, and drive reliability across GPU fleets.

You will develop Python tooling, implement security baselines, and optimize storage and Linux configurations for AI workloads. Relocation to San Francisco is offered and benefits include health/dental/vision.

Qualifications

  • 3+ years managing bare-metal and cloud-based server fleets at scale (100+ nodes).
  • Strong software engineering skills in Python; production tooling, not scripts.
  • Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd.

Responsibilities

  • Build and maintain Python fleet tracking system for servers: provisioning, health, RMAs, pricing, availability.
  • Develop tooling to automate provisioning, health checks, GPU diagnostics, recovery, and alerting.
  • Create metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk, network).
  • Leverage AI to automate alerting and recovery processes.
  • Implement OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning.
  • Manage and optimize storage systems for model weights, checkpoints, scratch space (NVMe, NFS, Lustre/GPFS).
  • Tune Linux for AI workloads: kernel, NUMA, CPU pinning, hugepages, I/O schedulers, NVIDIA driver stack.
  • Develop a suite of automated error detection and recovery processes.
  • Work with partners to solve technical issues.

Skills

Python
Linux systems
Infrastructure as code
System monitoring
Security hardening

Tools

Ansible
Terraform
cloud-init

Job description

fal is building a scalable generative AI infrastructure. We seek an hands-on engineer to manage thousands of servers, automate provisioning, health checks, and recovery, and drive reliability across GPU fleets.

You will develop Python tooling, implement security baselines, and optimize storage and Linux configurations for AI workloads. Relocation to San Francisco is offered and benefits include health/dental/vision.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra Engineer: GPU Fleet Automation
AI Infra Engineer: GPU Fleet Automation

Fal.ai Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Regular team events and offsites
Platform Software Engineer—AI Infra for GPU Fleet
Platform Software Engineer—AI Infra for GPU Fleet

fal - Features & Labels • United States

Remote
USD 150,000 - 210,000
Interesting and challenging work
Learning and growth opportunities
Visa sponsorship and relocation to San
Generative AI Infra Engineer (GPU Fleet)
Generative AI Infra Engineer (GPU Fleet)

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
Software Engineer, Infrastructure
Software Engineer, Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation to San Francisco
Health, dental, and vision insurance (
Regular team events
+1
Software Engineer, Infrastructure
Software Engineer, Infrastructure

The Consensus • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
Software Engineer, Platform
Software Engineer, Platform

fal - Features & Labels • United States

Remote
USD 150,000 - 210,000
Interesting and challenging work
Learning and growth opportunities
Visa sponsorship and relocation to San
Software Engineer, Platform
Software Engineer, Platform

Fal.ai Inc. • San Francisco (CA), Northern (KY)

On-site
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Regular team events and offsites
Senior AI Infra Engineer: High-Perf Kubernetes & GPUs
Senior AI Infra Engineer: High-Perf Kubernetes & GPUs

Fal.ai Inc. • Northern (KY)

Hybrid
USD 110,000 - 150,000
Senior GPU Cluster Infra Engineer | Remote
Senior GPU Cluster Infra Engineer | Remote

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3
Software Engineer, Distributed Systems
Software Engineer, Distributed Systems

fal - Features & Labels • San Francisco (CA)

On-site
USD 180,000 - 250,000
Health insurance
Relocation assistance
Vision insurance
+3