Member of Technical Staff, MLSys

Bake AI

San Mateo, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Bake AI in San Mateo, California, is hiring a Member of Technical Staff to build ML systems that enable sustained AI research. You will own performance-critical components across model serving, GPU execution, and distributed workloads, and you will push beyond a simple inference engine to improve correctness and efficiency.

You will work with researchers and infrastructure engineers, design robust systems, measure performance, and drive implementation from architecture through production, while

Qualifications

  • Independently designed and shipped a substantial ML systems component in production or a research setting.
  • Strong Python and systems programming with C++ or Rust and ability to modify an execution engine.
  • Hands-on GPU programming and profiling with CUDA or Triton and memory hierarchy understanding.
  • Experience with transformer execution, KV caches, and distributed LLM serving across multiple GPUs.

Responsibilities

  • Design and improve inference infrastructure for long-running agents and research workloads.
  • Extend or build inference runtimes when existing systems do not meet workload needs.
  • Evaluate techniques for latency, throughput, memory, and cost and implement trade-offs.
  • Profile full execution path from Python to GPU kernels and memory movement.
  • Build distributed execution paths for training, post-training, and agent rollouts.
  • Establish reproducible benchmarks and regression checks for correctness and performance.
  • Own designs and production follow-through with strong code review and collaboration.

Skills

Python
C++
Rust
GPU programming
CUDA
Triton
Distributed systems
Performance optimization

Tools

vLLM
SGLang
Profiling tools

Job description

We are hiring one Member of Technical Staff to build the ML systems that make sustained AI research possible. This is a full-time position based in San Mateo, California.

You will own performance-critical systems across model serving, GPU execution, and distributed research workloads. The work goes beyond deploying an inference engine: you will understand its internals, identify where its abstractions break down, and make changes that improve performance without compromising correctness or reliability.

We are looking for someone who can take an ambiguous systems problem from first principles through architecture, implementation, measurement, and deployment. You will work directly with researchers and infrastructure engineers, set a high technical standard, and remain deeply involved in the code.

What you’ll do
  • Design and improve inference infrastructure for long-running agents and research workloads, including scheduling, batching, model routing, and resource isolation.
  • Extend inference runtimes such as vLLM or SGLang, or build specialized components when existing systems do not meet the workload’s needs.
  • Evaluate and implement techniques such as prefill/decode disaggregation, speculative decoding, prefix reuse, and KV-cache management, with explicit trade-offs between latency, throughput, memory, and cost.
  • Profile and optimize the full execution path, from Python and host-side scheduling to GPU kernels, memory movement, and communication across devices and nodes.
  • Build distributed execution paths for training, post-training, and agent rollouts; connect research workflows to reliable compute and model-serving systems.
  • Establish reproducible benchmarks and regression checks for correctness, time to first token, inter-token latency, tail latency, throughput, and resource efficiency under realistic load.
  • Own technical designs and production follow-through: review code, diagnose difficult failures, document decisions, and turn one-off improvements into maintainable systems.
What you’ll bring
  • Evidence of having independently designed and shipped a substantial ML systems component used in production or by a research team. You can explain its architecture, failure modes, and the results you personally delivered.
  • Strong Python and systems programming skills in C++ or Rust, with the ability to work inside an execution engine rather than only configure its public APIs.
  • Hands-on GPU programming and profiling experience with CUDA or Triton. You understand memory hierarchy, bandwidth, synchronization, occupancy, and numerical correctness well enough to diagnose and implement meaningful optimizations.
  • A deep understanding of transformer execution, attention and KV caches, precision and quantization, and how model architecture changes compute and memory requirements.
  • Practical experience operating and improving LLM serving systems across multiple GPUs, including concurrency, backpressure, capacity planning, and failure recovery.
  • Strong distributed systems fundamentals, including communication/computation trade-offs, parallelism strategies, collective operations, and debugging failures that span processes or machines.
  • A rigorous approach to performance claims: representative workloads, controlled comparisons, profiling evidence, and checks that gains hold under realistic conditions.
  • The judgment to set technical direction, challenge assumptions, and carry a difficult project through with limited oversight. Clear writing, careful code review, and productive collaboration are part of the role.
Preferred experience
  • Significant contributions to an inference engine, GPU kernel library, compiler, or distributed training framework.
  • Experience with FSDP, DeepSpeed, Megatron, or related systems, especially online training or reinforcement learning workloads that coordinate generation and optimization.
  • Work on tensor, pipeline, or expert parallelism, high-speed interconnects, NCCL, or topology-aware execution.
  • Research publications in ML systems, computer architecture, compilers, or distributed systems, with substantial implementation responsibility.
  • Experience translating a research prototype into a reliable system that other teams depend on.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Physical Superintelligence • Boston (MA)

Hybrid
USD 140,000 - 210,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Member of Technical Staff, Inference Systems
Member of Technical Staff, Inference Systems

Confidential • California (MO)

On-site
USD 150,000 - 210,000
Member of Technical Staff (Performance Optimization)
Member of Technical Staff (Performance Optimization)

Fireworks AI • United States

On-site
USD 150,000 - 260,000
Member of Technical Staff, AI Infrastructure
Member of Technical Staff, AI Infrastructure

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 180,000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • New York (NY)

On-site
USD 220,000 - 485,000
AI Systems Engineer
AI Systems Engineer

Transluce • San Francisco (CA)

On-site
USD 350,000 - 600,000
Member of Technical Staff - ML Infra
Member of Technical Staff - ML Infra

Kindredventures • San Francisco (CA)

On-site
USD 160,000 - 220,000