Get more replies from employers
Send a job-specific resume in minutes.
Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control, queuing, and fairness across a global fleet. You will also build LV routing systems to dispatch workloads with latency awareness and predictive autoscaling, and explore KV caching for memory/compute trade-offs in LLM inference stacks.
You will contribute to deep observability, tracing every millisecond and preemptively addressing failures before customers
Sail is the foundation of useful, agentic AI. We are here to take a big swing at the most ambitious engineering challenge of our careers. Everyone working at Sail will become an expert; nothing less will do in our immensely competitive market.
Build the systems that make AI inference fast, reliable, and cost-efficient at global scale. You’ll design the control plane that schedules a huge queue of tokens over a diverse fleet of machines, spread all over the world.
We work out of a beautiful, sunny office in downtown San Francisco. All meals are on us (and actually great; SF is a food paradise and it would be a shame to eat only bowl slop). Everyone gets a Studio Display at their desk. We are serious about investing in anything that saves us time or energy. There are six different ways to make coffee or tea in the office. A friendly (hypoallergenic) black cat named Coco visits occasionally.