ML Systems Engineer: Inference & GPU-Driven Distributed Workloads

Bake AI

San Mateo, Northern (CA, KY)

Hybrid

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Bake AI in San Mateo, California, is hiring a Member of Technical Staff to build ML systems that enable sustained AI research. You will own performance-critical components across model serving, GPU execution, and distributed workloads, and you will push beyond a simple inference engine to improve correctness and efficiency.

You will work with researchers and infrastructure engineers, design robust systems, measure performance, and drive implementation from architecture through production, while

Qualifications

  • Independently designed and shipped a substantial ML systems component in production or a research setting.
  • Strong Python and systems programming with C++ or Rust and ability to modify an execution engine.
  • Hands-on GPU programming and profiling with CUDA or Triton and memory hierarchy understanding.
  • Experience with transformer execution, KV caches, and distributed LLM serving across multiple GPUs.

Responsibilities

  • Design and improve inference infrastructure for long-running agents and research workloads.
  • Extend or build inference runtimes when existing systems do not meet workload needs.
  • Evaluate techniques for latency, throughput, memory, and cost and implement trade-offs.
  • Profile full execution path from Python to GPU kernels and memory movement.
  • Build distributed execution paths for training, post-training, and agent rollouts.
  • Establish reproducible benchmarks and regression checks for correctness and performance.
  • Own designs and production follow-through with strong code review and collaboration.

Skills

Python
C++
Rust
GPU programming
CUDA
Triton
Distributed systems
Performance optimization

Tools

vLLM
SGLang
Profiling tools

Job description

Bake AI in San Mateo, California, is hiring a Member of Technical Staff to build ML systems that enable sustained AI research. You will own performance-critical components across model serving, GPU execution, and distributed workloads, and you will push beyond a simple inference engine to improve correctness and efficiency.

You will work with researchers and infrastructure engineers, design robust systems, measure performance, and drive implementation from architecture through production, while

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer — On-Site in Palo Alto, High-Impact
ML Systems Engineer — On-Site in Palo Alto, High-Impact

Recruiting From Scratch • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Competitive equity
Cutting-edge diffusion models
Direct collaboration with researchers
ML Systems Engineer: Scale Training & Inference
ML Systems Engineer: Scale Training & Inference

Doist • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive cash compensation
Startup equity
Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
GenAI ML Systems Engineer: Scalable Training & Inference
GenAI ML Systems Engineer: Scalable Training & Inference

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
Member of Technical Staff, MLSys
Member of Technical Staff, MLSys

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
ML Systems Engineer for RL & Inference Infrastructure
ML Systems Engineer for RL & Inference Infrastructure

Advanced Micro Devices • Santa Clara (CA)

Hybrid
USD 160,000 - 210,000
AMD benefits
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1