Senior ML Infra Engineer: Scalable AI Training Systems

Preference Model

Seattle (WA)

On-site

USD 180,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Vision insurance
Dental insurance
401K match
Lunch onsite
Weekly snacks

Job summary

Preference Model is seeking a Senior ML Infrastructure Engineer to build scalable systems powering post-training research on large language models. You will design compute, scheduling, and data infrastructure for in-house RL environments and maintain core ML framework primitives to accelerate experiments.

You will partner with Research Engineers to translate research needs into infrastructure requirements, ensuring reliability, low latency, and high throughput while scaling with cutting-edge AI

Qualifications

  • Strong software engineering fundamentals for production-grade infra.
  • Experience building ML/data-intensive infrastructure.
  • Familiarity with ML frameworks such as PyTorch or JAX.

Responsibilities

  • Design, build, and scale compute, scheduling, and data infrastructure for post-training research in in-house RL environments.
  • Develop core ML framework primitives and internal tooling for reproducible experimentation.
  • Build evaluation and benchmarking infra, monitoring, logging, and deployment systems to keep infra reliable as it scales.
  • Collaborate with Research Engineers to translate research needs into infrastructure requirements and ship rapidly.

Skills

Software engineering
ML frameworks (PyTorch/JAX)
Distributed systems
Cloud platforms (AWS/GCP)
Kubernetes
Data pipelines
LLM training/inference basics
Infra reliability

Tools

JVM/CI tooling

Job description

Preference Model is seeking a Senior ML Infrastructure Engineer to build scalable systems powering post-training research on large language models. You will design compute, scheduling, and data infrastructure for in-house RL environments and maintain core ML framework primitives to accelerate experiments.

You will partner with Research Engineers to translate research needs into infrastructure requirements, ensuring reliability, low latency, and high throughput while scaling with cutting-edge AI

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infrastructure Engineer — High-Throughput AI Research
Senior ML Infrastructure Engineer — High-Throughput AI Research

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Senior ML Systems Engineer: Scalable Training Frameworks
Senior ML Systems Engineer: Scalable Training Frameworks

Cohere • San Francisco (CA)

Hybrid
USD 150,000 - 180,000
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Senior Staff ML Engineer - Scalable LLM Infra
Senior Staff ML Engineer - Scalable LLM Infra

Moveworks • Mountain View (CA), Northern (KY)

Hybrid
USD 190,000 - 280,000
Senior ML Infra Engineer — Scalable AI for Neuroscience
Senior ML Infra Engineer — Scalable AI for Neuroscience

Stealth Neurotechnology Company • San Francisco (CA)

On-site
USD 180,000 - 230,000
Stock options
Comprehensive benefits
401(k) with matching
Staff Engineer, Scalable RL Infrastructure
Staff Engineer, Scalable RL Infrastructure

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior ML Systems Engineer: Training Infra
Senior ML Systems Engineer: Training Infra

Neura Market • San Francisco (CA)

On-site
USD 295,000 - 380,000
Relocation assistance
Senior ML Infra Engineer - Build Scalable AI for Legal
Senior ML Infra Engineer - Build Scalable AI for Legal

General Legal • New York (NY), San Francisco (CA)

Hybrid
USD 200,000 - 275,000
RL Systems Architect: Scalable AI Training & Infra
RL Systems Architect: Scalable AI Training & Infra

Bytedance • San Jose (CA)

On-site
USD 244,000 - 450,000