Software Engineer, Distributed Systems

Fal

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote work options for senior levels
Visa sponsorship and relocation
Health, dental, and vision insurance
Equity and competitive salary

Job summary

fal in San Francisco is seeking an experienced Software Engineer, Distributed Systems to build large-scale platform components in Python or Rust. You will own request routing, AI workload orchestration, scheduling, and GPU autoscaling for a rapidly growing service serving millions of users.

This full-time in-person role offers compelling compensation and visa sponsorship, with remote options considered for Senior and Staff levels.

Qualifications

  • 3+ years experience building distributed compute and orchestration platforms in Python or Rust.
  • Strong understanding of distributed systems fundamentals: consensus, scheduling, fault tolerance, capacity planning.
  • Deep understanding of computational complexity and memory allocation.
  • Track record of designing systems that scale under real production load.
  • Experience building and using observability to drive performance and reliability decisions.
  • Excellent communication and ability to drive technical decisions across teams.
  • Self‑starter who executes quickly, takes ownership, and constantly seeks improvement.

Responsibilities

  • Build our core Python/Rust platform: request routing, AI workload orchestration, scheduling, GPU autoscaling, large‑scale file storage, queueing, etc.
  • Produce forward designs for platform evolution as we scale to 100x current traffic and need to provide low latency across the world.
  • Leverage AI to an extreme level to automate the mundane parts of building complex but reliable systems.
  • Profile and tune low‑level CPU and memory performance.

Skills

Distributed systems
Python
Rust
Observability
Communication
Ownership

Tools

Async runtimes
Zero-copy
Memory-safe concurrency

Job description

Hi there! We are fal, and we are on a mission to build world’s first generative media platform for developers.

We built a serverless runtime for Python that is optimized to run large ML models on 1000s of GPUs efficiently. The applications built on our platform are currently serving millions of users around the world and our goal is to 1000x that over the next few years.

fal is an in‑person company based in San Francisco. Today, we are 80 people strong, and we are looking for team members who share our excitement about the fast‑moving nature of AI and can independently build world‑class infrastructure.

Software Engineer, Distributed Systems
Location

San Francisco, CA (willing to consider remote for Senior and Staff levels)

Employment Type

Full time

Department

fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high‑performance inference, orchestration, and observability come together to unlock new categories of AI‑native products.

About this role

You are an experienced software engineer who thrives on building large‑scale computing platforms. You have deep expertise in large scale distributed systems that deal with high complexity, a lot of traffic and data. You know how to achieve reliability and scale with minimum operational load.

Key responsibilities
  • Build our core Python/Rust platform: request routing, AI workload orchestration, scheduling, GPU autoscaling, large‑scale file storage, queueing, etc.
  • Produce forward designs for platform evolution as we scale to 100x current traffic and need to provide low latency across the world.
  • Leverage AI to an extreme level to automate the mundane parts of building complex but reliable systems.
  • Profile and tune low‑level CPU and memory performance.
Requirements
  • 3+ years experience building distributed compute and orchestration platforms in Python or Rust.
  • Strong understanding of distributed systems fundamentals: consensus, scheduling, fault tolerance, capacity planning.
  • Deep understanding of computational complexity and memory allocation.
  • Track record of designing systems that scale under real production load.
  • Experience building and using observability to drive performance and reliability decisions.
  • Excellent communication and ability to drive technical decisions across teams.
  • Self‑starter who executes quickly, takes ownership, and constantly seeks improvement.
Nice to have
  • Experience with AI/ML inference or training infrastructure.
  • Experience with high‑performance systems programming (async runtimes, zero‑copy, memory‑safe concurrency).
  • Background in building multi‑tenant compute platforms.
  • Understanding of networking fundamentals and performance characteristics.
  • Familiarity with GPU workload characteristics and scheduling constraints.
Compensation

$180,000-250,000 plus equity + benefits (This range is across all 3 levels Mid, Senior and Staff)

Benefits
  • Interesting and challenging work
  • Competitive salary and equity
  • We prefer to work in‑person but also offer remote work opportunities for exceptional candidates.
  • We offer visa sponsorship and will help you relocate to San Francisco.
  • Health, dental, and vision insurance (US)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Site Reliability
Software Engineer, Site Reliability

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Visa sponsorship
Competitive salary and equity
+1
Senior Software Engineer, Product
Senior Software Engineer, Product

Fal • San Francisco (CA)

Hybrid
USD 180,000 - 230,000
Relocation assistance
Health, dental, and vision insurance
Regular team events
+2
Software Engineer, Infrastructure
Software Engineer, Infrastructure

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 250,000
Relocation assistance
Health, dental, and vision insurance (
Team events & offsites
+1
Senior Software Engineer, Product
Senior Software Engineer, Product

The Consensus • San Francisco (CA)

On-site
USD 180,000 - 230,000
Relocation assistance to SF
Health, dental, and vision insurance (
Regular team events and offsite
Software Engineer, Infrastructure
Software Engineer, Infrastructure

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Health, dental and vision insurance (U
Regular team events and offsites
Senior Software Engineer, Data
Senior Software Engineer, Data

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 225,000
Relocation assistance
Health, dental and vision insurance (U
Team events and offsites
+1
Senior Software Engineer, Core Product Systems
Senior Software Engineer, Core Product Systems

fal • San Francisco (CA)

On-site
USD 190,000 - 230,000
Relocation assistance + visa Spons. to
Health, dental, and vision insurance
Regular team events and offsites
Staff Software Engineer, ML Performance & Systems
Staff Software Engineer, ML Performance & Systems

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Competitive salary and equity
Visa sponsorship
Health, dental, and vision insurance
+1
Founding GTM Recruiter
Founding GTM Recruiter

Fal • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 200,000
Health, dental, and vision insurance (
Visa sponsorship
Relocation assistance
+2
Software Engineer, Growth
Software Engineer, Growth

fal • San Francisco (CA)

On-site
USD 170,000 - 220,000
Relocation assistance to San Francisco
Health, dental, and vision insurance (
Team events & offsites
+2