An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Baseten powers mission-critical inference for AI companies, uniting research, flexible infrastructure, and developer tooling to deploy cutting-edge models into production. We host Model APIs and hosted endpoints, prioritizing performance, scalability, and reliability across the platform.
As a distributed systems engineer on the Inference Platform, you will own end-to-end projects—from architecture through deployment and monitoring—on Kubernetes, focusing on fast, reliable, and cost-efficient
Python Kubernetes Distributed Systems LLM Inference API Development Networking GPU Workloads Observability Capacity Planning Debugging Collaboration Communication
ABOUT BASETENBaseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.THE ROLEWe're looking for distributed systems engineers and product-minded generalists to build the distributed runtime that powers large-scale LLM inference on Baseten. Our inference platform empowers customers to deploy and operate cutting-edge models with industry-leading performance, scalability, and reliability. It also powers Model APIs, our hosted endpoints for the latest open-source models. You'll work across the stack, from the developer experience customers use to deploy models, through the libraries behind features like tool calling and reasoning, down to the systems that orchestrate deployments on Kubernetes and route traffic efficiently. Your job is to make sure every model on our platform is fast, reliable, and cost-efficient. You'll join a small, high-impact team at the intersection of distributed systems, model performance, infrastructure, and product, helping define how developers use AI models at scale. This role is ideal for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users. EXAMPLE INITIATIVESYou'll get to work on these types of projects on our Inference Platform:- The Baseten Inference Stack at NVIDIA Dynamo Day https://www.baseten.co/blog/nvidia-dynamo-day-baseten-inference-stack/- 2x faster inference with KV cache-aware routing https://www.baseten.co/blog/how-baseten-achieved-2x-faster-inference-with-nvidia-dynamo/- How Baseten multi-cloud capacity management (MCM) unifies deployments https://www.baseten.co/blog/how-baseten-multi-cloud-capacity-management-mcm-powers-cloud-self-hosted-and-hybr/#comparing-deployment-options-cloud-vs-self-hosted-vs-hybrid
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status. We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).