Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc

San Francisco (CA)

On-site

USD 300,000 - 550,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity compensation
401(k) matching
Health, dental and vision insurance
Unlimited PTO
Visa sponsorship
Relocation stipend
Small, focused team

Job summary

Magic AI, Inc. is seeking a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows.

You will work where model execution meets distributed infrastructure, influencing latency, throughput, and reliability of RL and training loops. You will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-cache strategies, batching, and long-context workloads while collaborating with

Qualifications

  • Strong software engineering and distributed systems fundamentals.
  • Experience building or operating large-scale inference or training systems.
  • Deep understanding of GPU execution constraints and memory trade-offs.
  • Experience debugging performance issues in production ML systems.
  • Ability to reason about system-level trade-offs between latency, throughput and cost.

Responsibilities

  • Design and scale high-performance inference serving systems.
  • Optimize KV-cache management, batching strategies, and scheduling.
  • Improve throughput and latency for long-context workloads.
  • Build and maintain distributed RL and post-training infrastructure.
  • Improve reliability of rollout, evaluation, and reward pipelines.
  • Automate fault detection and recovery for serving and RL systems.
  • Profile and eliminate performance bottlenecks across GPU, networking and storage layers.
  • Collaborate with Kernels and Research to align execution systems with model architecture.

Skills

Distributed systems
Production ML systems
GPU memory constraints
Performance debugging
Latency vs throughput
Ownership of production infra

Job description

Magic AI, Inc. is seeking a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows.

You will work where model execution meets distributed infrastructure, influencing latency, throughput, and reliability of RL and training loops. You will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-cache strategies, batching, and long-context workloads while collaborating with

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, RL Inference & Distributed Systems
Staff Engineer, RL Inference & Distributed Systems

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 230,000
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Infra Engineer: Scalable AI Training Systems
Senior ML Infra Engineer: Scalable AI Training Systems

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Staff AI Systems Engineer — Inference & RL
Staff AI Systems Engineer — Inference & RL

Together • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Remote ML Systems Engineer: Scalable AI Inference
Remote ML Systems Engineer: Scalable AI Inference

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Senior Staff ML Engineer - Scalable LLM Infra
Senior Staff ML Engineer - Scalable LLM Infra

Moveworks • Mountain View (CA), Northern (KY)

Hybrid
USD 190,000 - 280,000
Member of Technical Staff, Inference & RL Systems
Member of Technical Staff, Inference & RL Systems

Magic • San Francisco (CA)

On-site
USD 225,000 - 550,000
Equity compensation
401(k) with salary matching
Generous health, dental, and vision insurance
+2
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000