ML Infra Engineer — Scalable Training Systems

Monograph

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading tech company in San Francisco seeks a Machine Learning Engineer to build and maintain infrastructure for large-scale model training. In this hands-on role, you will design systems, work closely with researchers, and optimize training processes. Candidates should have strong software engineering skills and experience with JAX or PyTorch. Join a dynamic team at the forefront of machine learning and contribute to core training code and systems.

Qualifications

  • Strong experience building ML training infrastructure or internal platforms.
  • Hands-on experience with large-scale training in JAX or PyTorch.
  • Ability to optimize performance across the training stack.

Responsibilities

  • Design and maintain systems for large-scale model training.
  • Scale JAX-based training across TPU and GPU clusters.
  • Optimize memory usage, device utilization, and throughput.

Skills

Software engineering fundamentals
Large-scale training experience in JAX
Distributed training familiarity
Managing training workloads on cloud platforms
Debugging performance bottlenecks
Cross-functional communication

Education

Bachelor's degree in a related field

Tools

JAX
PyTorch
Kubernetes
SLURM
GCP TPU/GKE
AWS

Job description

A leading tech company in San Francisco seeks a Machine Learning Engineer to build and maintain infrastructure for large-scale model training. In this hands-on role, you will design systems, work closely with researchers, and optimize training processes. Candidates should have strong software engineering skills and experience with JAX or PyTorch. Join a dynamic team at the forefront of machine learning and contribute to core training code and systems.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infra Engineer: Build Scalable LLM Systems
Senior ML Infra Engineer: Build Scalable LLM Systems

ServiceNow • Mountain View (CA)

On-site
USD 130,000 - 180,000
ML Infra Engineer: Scale Training & Inference (Hybrid)
ML Infra Engineer: Scale Training & Inference (Hybrid)

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Infra Engineer
ML Infra Engineer

Monograph • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Infra Architect: Build Scalable Data & Training Pipelines
ML Infra Architect: Build Scalable Data & Training Pipelines

Humble Robotics • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior ML Infra Engineer for Scalable LLM Systems
Senior ML Infra Engineer for Scalable LLM Systems

Moveworks • Mountain View (CA)

On-site
USD 120,000 - 160,000
Senior ML Systems Engineer: Scalable Training Frameworks
Senior ML Systems Engineer: Scalable Training Frameworks

Cohere • San Francisco (CA)

Hybrid
USD 150,000 - 180,000
ML Infra Engineer: Scale GPU Training & Inference
ML Infra Engineer: Scale GPU Training & Inference

Reducto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Free lunch
Reimbursed transportation
+3
ML Infra Architect – On-Site in SF
ML Infra Architect – On-Site in SF

OP Recruiting • San Francisco (CA)

On-site
USD 180,000 - 260,000
On-site work
Health insurance
Wellness stipend
+3
ML Infrastructure Engineer for Scalable LLMs
ML Infrastructure Engineer for Scalable LLMs

ServiceNow • Mountain View (CA)

Hybrid
USD 130,000 - 160,000
Infrastructure Research Engineer: Scalable ML Training
Infrastructure Research Engineer: Scalable ML Training

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1