Staff Research Engineer — Large-Scale ML Pre-Training

magic.dev

San Francisco (CA)

On-site

USD 275,000 - 550,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
401(k) matching
Health insurance
Unlimited PTO
Visa sponsorship
Relocation stipend

Job summary

Magic is building safe AGI to accelerate humanity's progress. As a Research Engineer on the Pre-training Systems team, you will design and operate the distributed infrastructure that trains Magic’s long-context models at scale.

This role focuses on large-scale model training across GPU clusters, balancing performance, reliability and reproducibility while tackling memory pressure and cross-device communication in production ML systems.

Qualifications

  • Strong software engineering and distributed systems fundamentals.
  • Experience training large models in multi-node GPU environments.
  • Deep understanding of parallelism strategies and performance trade-offs.
  • Experience debugging cross-layer issues in production ML systems.
  • Strong ownership mindset and ability to operate critical infrastructure.
  • Track record of improving performance or reliability of large-scale systems.

Responsibilities

  • Scale distributed training across large GPU clusters.
  • Optimize communication patterns and gradient synchronization.
  • Improve checkpointing, fault tolerance, and job recovery systems.
  • Profile and eliminate performance bottlenecks across compute, networking, and storage.
  • Improve experiment reproducibility and orchestration workflows.
  • Increase hardware utilization and training throughput.
  • Collaborate with Kernels and Research to align model architecture with systems realities.

Skills

Software engineering
Distributed systems
Large-model training
Performance debugging

Job description

Magic is building safe AGI to accelerate humanity's progress. As a Research Engineer on the Pre-training Systems team, you will design and operate the distributed infrastructure that trains Magic’s long-context models at scale.

This role focuses on large-scale model training across GPU clusters, balancing performance, reliability and reproducibility while tackling memory pressure and cross-device communication in production ML systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Pre-training Systems
Member of Technical Staff, Pre-training Systems

magic.dev • San Francisco (CA)

On-site
USD 275,000 - 550,000
Equity
401(k) matching
Health insurance
+3
Staff Engineer, Inference & RL Systems — Scale ML
Staff Engineer, Inference & RL Systems — Scale ML

magic.dev • San Francisco (CA)

On-site
USD 275,000 - 550,000
Equity
401(k) matching
Health insurance
+4
Member of Technical Staff, Inference & RL Systems
Member of Technical Staff, Inference & RL Systems

magic.dev • San Francisco (CA)

On-site
USD 275,000 - 550,000
Equity
401(k) matching
Health insurance
+4
Research Engineer, RL Post-Training & Environments
Research Engineer, RL Post-Training & Environments

magic.dev • San Francisco (CA)

On-site
USD 275,000 - 550,000
Salary range 275k-550k
Equity compensation
401(k) with 6% match
+5
Member of Technical Staff, RL Research & Environments
Member of Technical Staff, RL Research & Environments

magic.dev • San Francisco (CA)

On-site
USD 275,000 - 550,000
Salary range 275k-550k
Equity compensation
401(k) with 6% match
+5
Member of Technical Staff, Inference & RL Systems
Member of Technical Staff, Inference & RL Systems

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Staff Research Engineer: Large-Scale AI Training
Staff Research Engineer: Large-Scale AI Training

Black Forest Labs Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 290,000
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO
Distributed ML Training Engineer - Scale GPUs, Unlimited PTO

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
AI Pre-Training Engineer — Massive GPU Scale
AI Pre-Training Engineer — Massive GPU Scale

OP Recruiting • Chicago (IL)

On-site
USD 160,000 - 260,000