Senior Distributed Systems Engineer - Scalable AI Platform

fal - Features & Labels

United States

Remote

USD 140.000 - 190.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Bekomme eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Interesting work
Learning opportunities
Team offsites

Zusammenfassung

fal is building a large-scale distributed platform for generative media, handling high traffic with reliable orchestration and observability. As a software engineer, you will contribute to the core Python and Rust platform, focusing on request routing, AI workload scheduling, and GPU autoscaling to support global production workloads.

You will design scalable systems, profile performance, and leverage AI to automate routine engineering tasks.

Qualifikationen

  • 5+ years experience building distributed compute and orchestration platforms in Python or Rust.
  • Strong understanding of distributed systems fundamentals: consensus, scheduling, fault tolerance, capacity planning.
  • Deep understanding of computational complexity and memory allocation.
  • Track record of designing systems that scale under real production load.
  • Experience building and using observability to drive performance and reliability decisions.
  • Excellent communication and ability to drive technical decisions across teams.
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement.

Aufgaben

  • Build our core Python/Rust platform: request routing, AI workload orchestration, scheduling, GPU autoscaling, large scale file storage, queueing, etc
  • Produce forward designs for platform evolution as we scale to 100x current traffic and need to provide low latency across the world
  • Leverage AI to an extreme level to automate the mundane parts of building complex but reliable systems
  • Profile and tune low level CPU and memory performance

Kenntnisse

Distributed compute
Python
Rust
Orchestration
Observability
Communication
Self-starter

Jobbeschreibung

fal is building a large-scale distributed platform for generative media, handling high traffic with reliable orchestration and observability. As a software engineer, you will contribute to the core Python and Rust platform, focusing on request routing, AI workload scheduling, and GPU autoscaling to support global production workloads.

You will design scalable systems, profile performance, and leverage AI to automate routine engineering tasks.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Software Engineer, Distributed Systems
Software Engineer, Distributed Systems

fal - Features & Labels • USA

Remote
USD 140.000 - 190.000
Interesting work
Learning opportunities
Team offsites
GPU Infra Engineer for Scalable AI Platform
GPU Infra Engineer for Scalable AI Platform

Kindredventures • San Francisco (CA)

Vor Ort
USD 180.000 - 250.000
Relocation to San Francisco
Health, dental, and vision insurance (
Regular team events
+1
Senior Full-Stack Engineer for Scalable AI Platform (SF)
Senior Full-Stack Engineer for Scalable AI Platform (SF)

The Consensus • San Francisco (CA)

Vor Ort
USD 180.000 - 230.000
Relocation assistance to SF
Health, dental, and vision insurance (
Regular team events and offsite
Senior Network Infrastructure Engineer: Scale & Performance
Senior Network Infrastructure Engineer: Scale & Performance

Kindredventures • USA

Vor Ort
USD 180.000 - 240.000
Senior Security Engineer: AI Infrastructure & GPU Compute
Senior Security Engineer: AI Infrastructure & GPU Compute

The Consensus • San Francisco (CA)

Vor Ort
USD 180.000 - 260.000
Equity
Health plan
Dental & Vision
Senior AI Infra Engineer: High-Perf Kubernetes & GPUs
Senior AI Infra Engineer: High-Perf Kubernetes & GPUs

Fal.ai Inc. • Northern (KY)

Hybrid
USD 110.000 - 150.000
Software Engineer, Site Reliability
Software Engineer, Site Reliability

fal - Features & Labels • USA

Remote
USD 120.000 - 180.000
Senior Staff Infra Engineer — AI Compute & GPU Clusters
Senior Staff Infra Engineer — AI Compute & GPU Clusters

Fal • USA

Remote
USD 180.000 - 250.000
Visa sponsorship
Relocation to San Francisco
Health insurance
Generative AI Infra Engineer (GPU Fleet)
Generative AI Infra Engineer (GPU Fleet)

The Consensus • San Francisco (CA)

Hybrid
USD 180.000 - 250.000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Fal.ai Inc. • Northern (KY)

Hybrid
USD 110.000 - 150.000