Senior Staff Infra Engineer — AI Compute & GPU Clusters

Fal

United States

Remote

USD 180,000 - 250,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Visa sponsorship
Relocation to San Francisco
Health insurance

Job summary

fal is building the high-performance compute environments for customers, spanning bare-metal servers, GPU-passthrough VMs, Kubernetes and Slurm clusters, and distributed storage. You will shape Linux images, provisioning, and cluster networking to deliver reliable, repeatable environments.

You will drive automation across provisioning, upgrades, and recovery, leveraging AI to accelerate every facet of infrastructure delivery and operations while collaborating with customers and internal teams.

Qualifications

  • Experience designing and operating production Linux infrastructure.
  • Proficient with Kubernetes on bare metal and container runtimes.
  • Familiarity with NVIDIA GPU drivers, GPU telemetry, and GPU acceleration tooling.
  • Strong networking fundamentals and distributed storage experience.

Responsibilities

  • Design, automate, validate, and deliver the complete lifecycle of customer compute environments—from provisioning through upgrades, recovery, and decommissioning.
  • Use AI to automate and accelerate infrastructure delivery and operations.
  • Provision dedicated Kubernetes and Slurm clusters for customer workloads.
  • Build and maintain Linux images and automated OS provisioning workflows.
  • Operate the NVIDIA GPU stack and related tooling.
  • Design Kubernetes and data-center networking and configure distributed storage for high-performance workloads.
  • Build monitoring, alerting, diagnostics, and automated recovery for customer environments.
  • Develop reusable tooling, standards, documentation, and runbooks.
  • Collaborate with customers and internal teams to translate workload requirements into infrastructure designs.

Skills

Strong communication
Ownership
Problem solving
Team collaboration

Tools

Kubernetes
Slurm
Linux
NVIDIA GPU stack
Cilium/Calico networking
Ceph/Lustre storage
GPU Operator
Docker/Containerd

Job description

fal is building the high-performance compute environments for customers, spanning bare-metal servers, GPU-passthrough VMs, Kubernetes and Slurm clusters, and distributed storage. You will shape Linux images, provisioning, and cluster networking to deliver reliable, repeatable environments.

You will drive automation across provisioning, upgrades, and recovery, leveraging AI to accelerate every facet of infrastructure delivery and operations while collaborating with customers and internal teams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer: High-Perf Kubernetes & GPUs
Senior AI Infra Engineer: High-Perf Kubernetes & GPUs

Fal.ai Inc. • Northern (KY)

Hybrid
USD 110,000 - 150,000
Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Fal.ai Inc. • Northern (KY)

Hybrid
USD 110,000 - 150,000
Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Kindredventures • United States

On-site
USD 140,000 - 190,000
GPU Infra Engineer for Scalable AI Platform
GPU Infra Engineer for Scalable AI Platform

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation to San Francisco
Health, dental, and vision insurance (
Regular team events
+1
Senior Distributed Systems Engineer - Scalable AI Platform
Senior Distributed Systems Engineer - Scalable AI Platform

fal - Features & Labels • United States

Remote
USD 140,000 - 190,000
Interesting work
Learning opportunities
Team offsites
Senior Network Infrastructure Engineer: Scale & Performance
Senior Network Infrastructure Engineer: Scale & Performance

Kindredventures • United States

On-site
USD 180,000 - 240,000
Senior Network Infra Engineer: AI-Driven Connectivity
Senior Network Infra Engineer: AI-Driven Connectivity

fal - Features & Labels • United States

Remote
USD 180,000 - 240,000
Visa sponsorship
Relocation to San Francisco
Health, dental, and vision insurance (
Senior Network Engineer - AI-Driven Infrastructure
Senior Network Engineer - AI-Driven Infrastructure

Fal.ai Inc. • Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior GPU & Cloud Infrastructure Vendor Manager
Senior GPU & Cloud Infrastructure Vendor Manager

fal • San Francisco (CA)

On-site
USD 160,000 - 200,000
Senior GPU Cluster Infra Engineer | Remote
Senior GPU Cluster Infra Engineer | Remote

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3