Stand out for this role — generate a tailored resume and cover letter in about a minute.
SkyPilot is seeking an engineer to own the GPU and ML-systems layer that powers frontier AI workloads. You will shape accelerator scheduling, utilization, and the serving path across clouds and Kubernetes, delivering a fast, cost-efficient AI stack.
You will deepen integrations with vLLM, PyTorch, Slime and related frameworks while enabling scalable training, inference, and multi-cluster serving. You’ll work with a team moving cutting-edge AI compute forward at scale.
SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single \"AI supercomputer.\"
SkyPilot (10k+ GitHub stars, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.
SkyPilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on: accelerator scheduling and utilization, the serving path, and the integrations that make SkyPilot the fastest, most cost-efficient place to run demanding AI workloads. A few points of GPU utilization here can save a team millions in compute and days on every training run.
Location: San Mateo, CA. Remote will be considered for exceptional candidates.