Member of Technical Staff, Pre-training Systems

magic.dev

San Francisco (CA)

On-site

USD 275,000 - 550,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
401(k) matching
Health insurance
Unlimited PTO
Visa sponsorship
Relocation stipend

Job summary

Magic is building safe AGI to accelerate humanity's progress. As a Research Engineer on the Pre-training Systems team, you will design and operate the distributed infrastructure that trains Magic’s long-context models at scale.

This role focuses on large-scale model training across GPU clusters, balancing performance, reliability and reproducibility while tackling memory pressure and cross-device communication in production ML systems.

Qualifications

  • Strong software engineering and distributed systems fundamentals.
  • Experience training large models in multi-node GPU environments.
  • Deep understanding of parallelism strategies and performance trade-offs.
  • Experience debugging cross-layer issues in production ML systems.
  • Strong ownership mindset and ability to operate critical infrastructure.
  • Track record of improving performance or reliability of large-scale systems.

Responsibilities

  • Scale distributed training across large GPU clusters.
  • Optimize communication patterns and gradient synchronization.
  • Improve checkpointing, fault tolerance, and job recovery systems.
  • Profile and eliminate performance bottlenecks across compute, networking, and storage.
  • Improve experiment reproducibility and orchestration workflows.
  • Increase hardware utilization and training throughput.
  • Collaborate with Kernels and Research to align model architecture with systems realities.

Skills

Software engineering
Distributed systems
Large-model training
Performance debugging

Job description

Magic’s mission is to build safe AGI that accelerates humanity’s progress on the world’s most important problems. We believe the most promising path to safe AGI lies in automating research and code generation to improve models and solve alignment more reliably than humans can alone. Our approach combines frontier-scale pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve this goal.

About the role

As a Research Engineer on the Pre-training Systems team, you will design and operate the distributed infrastructure that trains Magic’s long-context models at scale.

This role focuses on large-scale model training across massive GPU clusters. You will work at the boundary between deep learning and distributed systems, ensuring that training runs are performant, reliable, and reproducible under extreme scale.

Magic’s long-context models create non-trivial systems challenges: sustained memory pressure, communication overhead across thousands of devices, long-running jobs that must survive failures, and efficient sequence packing under hardware constraints. You will own the systems that make large-scale pre-training stable and fast.

What you’ll work on
  • Scale distributed training across large GPU clusters (data, tensor, pipeline parallelism)

  • Optimize communication patterns and gradient synchronization

  • Improve checkpointing, fault tolerance, and job recovery systems

  • Profile and eliminate performance bottlenecks across compute, networking, and storage

  • Improve experiment reproducibility and orchestration workflows

  • Increase hardware utilization and training throughput

  • Collaborate with Kernels and Research to align model architecture with systems realities

What we’re looking for
  • Strong software engineering and distributed systems fundamentals

  • Experience training large models in multi-node GPU environments

  • Deep understanding of parallelism strategies and performance trade-offs

  • Experience debugging cross-layer issues in production ML systems

  • Strong ownership mindset and ability to operate critical infrastructure

  • Track record of improving performance or reliability of large-scale systems

Our culture
  • Integrity. Words and actions should be aligned

  • Hands-on. At Magic, everyone is building

  • Teamwork. We move as one team, not N individuals

  • Focus. Safely deploy AGI. Everything else is noise

  • Quality. Magic should feel like magic

Magic strives to be the place where high-potential individuals can do their best work. We value quick learning and grit just as much as skill and experience.

Compensation, benefits, and perks (US):
  • Annual salary range: $275K - $550K

  • Equity is a significant part of total compensation, in addition to salary

  • 401(k) plan with 6% salary matching

  • Generous health, dental and vision insurance for you and your dependents

  • Unlimited paid time off

  • Visa sponsorship and relocation stipend to bring you to SF, if possible

  • A small, fast-paced, highly focused team

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Inference & RL Systems
Member of Technical Staff, Inference & RL Systems

magic.dev • San Francisco (CA)

On-site
USD 275,000 - 550,000
Equity
401(k) matching
Health insurance
+4
Member of Technical Staff, Inference & RL Systems
Member of Technical Staff, Inference & RL Systems

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Member of Technical Staff, RL Research & Environments
Member of Technical Staff, RL Research & Environments

magic.dev • San Francisco (CA)

On-site
USD 275,000 - 550,000
Salary range 275k-550k
Equity compensation
401(k) with 6% match
+5
Member of Technical Staff, Security Engineer
Member of Technical Staff, Security Engineer

magic.dev • San Francisco (CA)

On-site
USD 225,000 - 550,000
Visa sponsorship
Relocation stipend to SF
Equity
+4
Head of IT
Head of IT

magic.dev • San Francisco (CA)

On-site
USD 200,000 - 350,000
Equity
Health, dental and vision insurance
Unlimited PTO
+2
Member of Technical Staff, Security Engineer
Member of Technical Staff, Security Engineer

Magic • San Francisco (CA)

On-site
USD 225,000 - 550,000
Equity
401(k) matching
Health, dental, and vision insurance
+3
Member of Technical Staff, Inference & RL Systems
Member of Technical Staff, Inference & RL Systems

Magic • San Francisco (CA)

On-site
USD 225,000 - 550,000
Equity compensation
401(k) with salary matching
Generous health, dental, and vision insurance
+2
Member of Technical Staff, Kernels
Member of Technical Staff, Kernels

Magic • United States

On-site
USD 225,000 - 550,000
401(k) plan with 6% salary matching
Generous health, dental and vision insurance
Unlimited paid time off
+1
Staff Research Engineer — Large-Scale ML Pre-Training
Staff Research Engineer — Large-Scale ML Pre-Training

magic.dev • San Francisco (CA)

On-site
USD 275,000 - 550,000
Equity
401(k) matching
Health insurance
+3
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

On-site
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1