Get more replies from employers
Send a job-specific resume in minutes.
Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control, queuing, and fairness across a global fleet. You will also build LV routing systems to dispatch workloads with latency awareness and predictive autoscaling, and explore KV caching for memory/compute trade-offs in LLM inference stacks.
You will contribute to deep observability, tracing every millisecond and preemptively addressing failures before customers
Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control, queuing, and fairness across a global fleet. You will also build LV routing systems to dispatch workloads with latency awareness and predictive autoscaling, and explore KV caching for memory/compute trade-offs in LLM inference stacks.
You will contribute to deep observability, tracing every millisecond and preemptively addressing failures before customers