Staff Engineer — AI Inference & Distributed Systems

Sail

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control, queuing, and fairness across a global fleet. You will also build LV routing systems to dispatch workloads with latency awareness and predictive autoscaling, and explore KV caching for memory/compute trade-offs in LLM inference stacks.

You will contribute to deep observability, tracing every millisecond and preemptively addressing failures before customers

Qualifications

  • Strong distributed systems fundamentals including concurrency, networking, databases, and performance engineering.

Responsibilities

  • Design and implement high-performance schedulers (admission control, queuing, priority, fairness, preemption, bin packing).
  • Build global routing and traffic management (latency-aware dispatch, predictive autoscaling, failover strategies).
  • LLM-specific routing optimizations, e.g. KV caching across GPU RAM, CPU RAM, NVMe flash.

Skills

Distributed systems
Testing & edge cases
Plan & docs

Job description

Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control, queuing, and fairness across a global fleet. You will also build LV routing systems to dispatch workloads with latency awareness and predictive autoscaling, and explore KV caching for memory/compute trade-offs in LLM inference stacks.

You will contribute to deep observability, tracing every millisecond and preemptively addressing failures before customers

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, AI Inference & Distributed Systems
Staff Engineer, AI Inference & Distributed Systems

Sail Research • San Francisco (CA)

On-site
USD 120,000 - 160,000
Free meals
Studio Display at desk
Friendly office environment with a cat
Member of Technical Staff - Distributed Systems
Member of Technical Staff - Distributed Systems

Sail • San Francisco (CA)

On-site
USD 150,000 - 190,000
Member of Technical Staff - Distributed Systems
Member of Technical Staff - Distributed Systems

Sail Research • San Francisco (CA)

On-site
USD 120,000 - 160,000
Free meals
Studio Display at desk
Friendly office environment with a cat
Staff AI Inference Kernel Engineer
Staff AI Inference Kernel Engineer

Sail • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff AI Systems Engineer — Inference & RL
Staff AI Systems Engineer — Inference & RL

Together • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Full Stack Engineer, Fleet Scheduling
Full Stack Engineer, Fleet Scheduling

Slope • San Francisco (CA)

On-site
USD 325,000 - 490,000
Staff Engineer, Distributed AI Inference Systems (Equity)
Staff Engineer, Distributed AI Inference Systems (Equity)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000
Equity
Staff Distributed Systems Engineer - AI Infrastructure
Staff Distributed Systems Engineer - AI Infrastructure

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Staff Software Engineer AI Inference Infra Orchestrator
Staff Software Engineer AI Inference Infra Orchestrator

Together AI • San Francisco (CA)

On-site
USD 240,000 - 280,000
Equity
Health insurance
Competitive benefits
Staff Engineer - Customer-Facing AI Inference Infra
Staff Engineer - Customer-Facing AI Inference Infra

Simplify • San Francisco (CA)

On-site
USD 200,000 - 300,000
Housing stipend
Uber/Waymo rides