Platform Software Engineer—AI Infra for GPU Fleet

fal - Features & Labels

United States

Remote

USD 150,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Interesting and challenging work
Learning and growth opportunities
Visa sponsorship and relocation to San

Job summary

fal - Features & Labels is seeking a hands-on engineer to manage a large pool of GPU servers. You will build Python-based fleet tooling, automate provisioning and health checks, and create dashboards for hardware health across thousands of servers.

You’ll harden OS security and optimize storage and GPU driver stacks, collaborating with partners to resolve complex infra issues. The role requires strong Python software engineering, deep Linux knowledge, and IaC experience.

Qualifications

  • 3+ years experience managing bare-metal and cloud based server fleets at scale (100+ nodes).
  • Strong software engineering skills in Python; production tooling mindset.
  • Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd.
  • IaC experience with Ansible, Terraform, cloud-init.
  • Solid understanding of storage technologies and hardware diagnostics.
  • Experience building internal tools or dashboards for infrastructure visibility.
  • Excellent communication and ability to drive technical decisions across teams.
  • Self-starter who executes quickly, takes ownership and seeks improvement.

Responsibilities

  • Build and maintain Python fleet tracking system managing server lifecycles, pricing, availability, and RMAs.
  • Develop server management tooling for provisioning, health checks, GPU diagnostics, recovery and alerting.
  • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet.
  • Leverage AI to automate alerting and recovery at scale.
  • Implement OS security baselines, SELinux/AppArmor, SSH key management, vulnerability scanning.
  • Manage distributed storage for model weights, checkpoints, and scratch space.
  • Tune Linux systems for AI workloads including kernel parameters and GPU driver stack.
  • Develop automated error detection and recovery processes and collaborate with partners.

Skills

Python
Linux systems
Communication
Self-starter
InfrastructureAutomation
Tooling development
Troubleshooting hardware
System performance

Tools

Ansible
Terraform
cloud-init
NVIDIA DCGM
libvirt/QEMU

Job description

fal - Features & Labels is seeking a hands-on engineer to manage a large pool of GPU servers. You will build Python-based fleet tooling, automate provisioning and health checks, and create dashboards for hardware health across thousands of servers.

You’ll harden OS security and optimize storage and GPU driver stacks, collaborating with partners to resolve complex infra issues. The role requires strong Python software engineering, deep Linux knowledge, and IaC experience.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Infra Engineer for Scalable AI Platform
GPU Infra Engineer for Scalable AI Platform

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation to San Francisco
Health, dental, and vision insurance (
Regular team events
+1
Generative AI Infra Engineer (GPU Fleet)
Generative AI Infra Engineer (GPU Fleet)

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
AI Infra Engineer: GPU Fleet Automation
AI Infra Engineer: GPU Fleet Automation

Fal.ai Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Regular team events and offsites
Staff Software Engineer - GPU Fleet & Cloud Infra
Staff Software Engineer - GPU Fleet & Cloud Infra

Cloudjobs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Restricted Stock Units
Health insurance
Vision insurance
+15
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery
Autonomous AI Infrastructure Engineer: GPU Fleet Mastery

Together • San Francisco (CA)

On-site
USD 190,000 - 270,000
Health insurance
Startup equity
Competitive benefits
GPU Fleet Orchestrator for AI Infra
GPU Fleet Orchestrator for AI Infra

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 150,000 - 210,000
GPU Fleet Orchestrator for AI Infrastructure
GPU Fleet Orchestrator for AI Infrastructure

AMD • San Jose (CA)

Hybrid
USD 180,000 - 240,000
AI Infrastructure Engineer — Automate GPU Fleets (Hybrid)
AI Infrastructure Engineer — Automate GPU Fleets (Hybrid)

WeHireYou • New Amsterdam (IN)

Hybrid
USD 113,000 - 181,000
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes
GPU Fleet Infra Engineer — Scale, Automation & Kubernetes

OpenAI • New York (NY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

Cloudjobs • New York (NY)

On-site
USD 120,000 - 160,000