Senior Research Engineer, Model Training / Pretraining

Intelix.AI

Seattle (WA)

On-site

USD 147,000 - 220,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Intelix.AI is seeking a Senior Research Engineer on the model scaling team, reporting to the Director of AI Research. The role covers model training and the supporting infrastructure, with training runs on 600 B300 and 1,000 H100 GPUs in a hybrid Seattle environment.

Responsibilities include building the training stack for distributed training, running scaling experiments, and specializing in CUDA kernels, Mixture-of-Experts, or RL post-training, plus releasing models and code publicly.

Qualifications

  • 6+ years of software engineering experience.
  • 4+ years of machine-learning infrastructure experience.
  • Experience training models from scratch (pretraining) is required; retrieval-augmented generation or fine-tuning only does not meet this.

Responsibilities

  • Build and maintain the training stack for distributed model training across GPU clusters.
  • Run scaling experiments covering model architecture, data, and optimisation.
  • Specialise in one area: CUDA kernels, Mixture-of-Experts, or reinforcement-learning post-training (GRPO, PPO, RLVR).
  • Release models, code, and associated research publicly.

Skills

Python
PyTorch
CUDA
Mixture-of-Experts
Reinforcement Learning

Tools

JAX

Job description

The position is a Senior Research Engineer on the model scaling team, reporting to a Director of AI Research. The work covers model training and the supporting infrastructure. Training runs use approximately 600 B300 and 1,000 H100 GPUs.

Responsibilities
  • Build and maintain the training stack for distributed model training across GPU clusters.
  • Run scaling experiments covering model architecture, data, and optimisation.
  • Work in one specialist area: GPU kernel development (CUDA), Mixture-of-Experts architectures, or reinforcement-learning post-training (GRPO, PPO, RLVR).
  • Release models, code, and associated research publicly.
Requirements
  • 6+ years of software engineering experience.
  • 4+ years of machine-learning infrastructure experience.
  • Python and PyTorch.
  • Experience training models from scratch (pretraining). Experience limited to retrieval-augmented generation or application-level fine-tuning does not meet this requirement.
  • Depth in at least one of: CUDA kernels, Mixture-of-Experts, or reinforcement learning (GRPO, PPO, RLVR).
Also considered
  • JAX.
  • Open-source contributions.
  • Published research.
  • Experience with GPU clusters at the scale described above.
  • Hiring-manager screen — 25 minutes.
  • Machine-learning coding interviews — two sessions, 45 minutes each.
  • Machine-learning system-design interview — 45 minutes.
  • Technical deep-dive with the Director of AI Research — 60 minutes.
  • Final interview — 30 minutes.
  • Offer.
  • Base salary: $147,000–$220,000 + strong bonus programmes in addition to base salary.
  • Hybrid in Seattle, Washington. On-site presence is required. Fully remote is not available.
  • Visa sponsorship is available.
Application

This search is managed by Intelix.AI.

Contact: Hasan Mohammad — hasan@intelix.ai.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Model Training & Scaling Engineer (Hybrid)
Senior Model Training & Scaling Engineer (Hybrid)

Intelix.AI • Seattle (WA)

On-site
USD 147,000 - 220,000
Member of Technical Staff: Training Infrastructure
Member of Technical Staff: Training Infrastructure

Wintermeyer Ventures • San Francisco (CA)

On-site
USD 200,000 - 375,000
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

On-site
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Research Engineer - Distributed Training
Research Engineer - Distributed Training

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 350,000
Helix AI Engineer, Training Performance
Helix AI Engineer, Training Performance

figure.ai • San Jose (CA), Northern (KY)

On-site
USD 200,000 - 400,000
Senior Solutions Architect, Robotics Foundation Model Training
Senior Solutions Architect, Robotics Foundation Model Training

NVIDIA • California (MO)

On-site
USD 152,000 - 242,000
Equity
Senior Solutions Architect, Robotics Foundation Model Training
Senior Solutions Architect, Robotics Foundation Model Training

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000