Stand out for this role — generate a tailored resume and cover letter in about a minute.
Baseten in San Francisco seeks an Engineering Manager for the Inference Performance team. You will lead engineers focused on making AI workloads faster on GPUs, including kernels, scheduling, batching, and KV-cache management.
This role blends technical depth with people leadership to deliver high-impact optimizations. You will define the team’s direction, hire and mentor engineers, and collaborate with Infrastructure, Model APIs and customer teams to ship performance wins and support new model
GPU Optimization Inference Performance Engineering CUDA TensorRT PyTorch Quantization Speculative Decoding KV-cache Management Batching Technical Leadership Team Management Performance Reviews
ABOUT BASETENBaseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.THE ROLEWe're looking for an Engineering Manager to lead part of our Inference Performance team. This team makes the world's most demanding AI workloads run faster and more efficiently on GPUs. You'll manage and grow a team of inference performance engineers working across the inference engine and runtime: kernels, scheduling, batching, KV-cache management, speculative decoding and prefill/decode disaggregation. This is a hands‑on technical leadership role. You'll set direction, unblock hard problems and earn the team's trust by going deep on GPU performance, while also hiring, developing and supporting the people doing the work. Your team's output directly affects how fast our customers' models run and how efficiently we serve them. The team is scaling quickly, so you'll help shape how it is structured as it grows.
Your team will work on these types of projects as part of our Inference Runtime team:
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status. We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).