Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Baseten is seeking an Engineering Manager to lead part of the Inference Performance team, driving faster, more efficient AI workloads on GPUs. You will manage and grow a team across inference engine and runtime, setting direction, hiring, and supporting the people doing the work.
You’ll shape how the team is structured as it scales, balance customer needs with platform investments and help bring new model architectures to production quickly.
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.
We're looking for an Engineering Manager to lead part of our Inference Performance team. This team makes the world's most demanding AI workloads run faster and more efficiently on GPUs. You'll manage and grow a team of inference performance engineers working across the inference engine and runtime: kernels, scheduling, batching, KV-cache management, speculative decoding and prefill/decode disaggregation. This is a hands-on technical leadership role. You'll set direction, unblock hard problems and earn the team's trust by going deep on GPU performance, while also hiring, developing and supporting the people doing the work. Your team's output directly affects how fast our customers' models run and how efficiently we serve them. The team is scaling quickly, so you'll help shape how it is structured as it grows.
Your team will work on these types of projects as part of our Inference Runtime team: