Staff Engineer, Inference & RL Systems — Scale ML

magic.dev

San Francisco (CA)

On-site

USD 275,000 - 550,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
401(k) matching
Health insurance
Unlimited PTO
Visa sponsorship
Relocation stipend
Small, focused team

Job summary

Magic is building safe AGI and operates distributed serving and RL infrastructure. As a Research Engineer on the Inference &RL Systems team, you will design and operate the systems that serve our models in production and power large-scale post-training workflows.

You will own the infrastructure that makes production inference fast and reliable, addressing KV-cache scaling, long-context workloads, and throughput under real-world workloads while collaborating with Kernels and Research.

Qualifications

  • Strong software engineering and distributed systems fundamentals.
  • Experience building or operating large-scale inference or training systems.
  • Deep understanding of GPU execution constraints and memory trade-offs.
  • Experience debugging performance issues in production ML systems.
  • Ability to reason about system-level trade-offs between latency, throughput, and cost.
  • Track record of owning critical production infrastructure.

Responsibilities

  • Design and scale high-performance inference serving systems
  • Optimize KV-cache management, batching strategies, and scheduling
  • Improve throughput and latency for long-context workloads
  • Build and maintain distributed RL and post-training infrastructure
  • Improve reliability of rollout, evaluation, and reward pipelines
  • Automate fault detection and recovery for serving and RL systems
  • Profile and eliminate performance bottlenecks across GPU, networking, and storage layers
  • Collaborate with Kernels and Research to align execution systems with model architecture

Skills

Distributed systems
Inference systems
GPU memory constraints
Performance debugging
System trade-offs
Production infra ownership

Job description

Magic is building safe AGI and operates distributed serving and RL infrastructure. As a Research Engineer on the Inference &RL Systems team, you will design and operate the systems that serve our models in production and power large-scale post-training workflows.

You will own the infrastructure that makes production inference fast and reliable, addressing KV-cache scaling, long-context workloads, and throughput under real-world workloads while collaborating with Kernels and Research.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Staff Research Engineer — Large-Scale ML Pre-Training
Staff Research Engineer — Large-Scale ML Pre-Training

magic.dev • San Francisco (CA)

On-site
USD 275,000 - 550,000
Equity
401(k) matching
Health insurance
+3
Member of Technical Staff, Inference & RL Systems
Member of Technical Staff, Inference & RL Systems

magic.dev • San Francisco (CA)

On-site
USD 275,000 - 550,000
Equity
401(k) matching
Health insurance
+4
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Member of Technical Staff, Inference & RL Systems
Member of Technical Staff, Inference & RL Systems

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
ML Infra Engineer: Scale RL Pipelines & Inference
ML Infra Engineer: Scale RL Pipelines & Inference

Moonfire • Paris (TX)

On-site
USD 115,000 - 173,000
Equity
Flexible time off
Relocation package
+4
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior ML Platform Engineer — Scale Research ML Infra
Senior ML Platform Engineer — Scale Research ML Infra

techire ai • San Francisco (CA)

On-site
USD 270,000 - 330,000
Stock options
Staff Engineer - Large-Scale GPU Inference & RL Infra
Staff Engineer - Large-Scale GPU Inference & RL Infra

reflectionai • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 320,000
Top-tier compensation
Stock options
Health & wellness
+5
Senior ML Inference Engineer — Scale Production APIs
Senior ML Inference Engineer — Scale Production APIs

AssemblyAI, Inc. • New York (NY)

On-site
USD 190,000 - 225,000