ML Systems Engineer: Scale Training & Inference

Doist

San Francisco (CA)

On-site

USD 180,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive cash compensation
Startup equity

Job summary

Engram is seeking an ML Systems Engineer to design, optimize, and scale training and inference workloads across GPUs and distributed systems. You’ll collaborate with researchers and performance engineers to bridge cutting‑edge AI research with production systems.

The role requires a Bachelor’s degree and 5+ years of experience in ML or training/inference, contributing to personalization features, memory retrieval, and low‑latency serving in a fast‑paced startup environment based in San Francisco.

Qualifications

  • 5+ years of experience with training or inference systems.
  • Strong engineering fundamentals and ability to ship high‑quality code in a fast‑paced environment.
  • Deep understanding of ML frameworks, GPUs, distributed systems, and infrastructure.
  • Experience with open‑source ML or systems projects is a plus.

Responsibilities

  • Design and optimize training and inference workloads for personalization and continual learning APIs.
  • Collaborate with researchers to turn prototypes into scalable production systems.
  • Improve serving paths for per‑user state and low latency requirements.
  • Work on distributed training across GPUs and nodes.

Skills

Training/inference systems
ML frameworks PyTorch/JAX
GPU computing & distributed systems
Open-source ML experience

Education

Bachelor’s degree or equivalent experience in CS/Engineering

Tools

PyTorch
JAX

Job description

Engram is seeking an ML Systems Engineer to design, optimize, and scale training and inference workloads across GPUs and distributed systems. You’ll collaborate with researchers and performance engineers to bridge cutting‑edge AI research with production systems.

The role requires a Bachelor’s degree and 5+ years of experience in ML or training/inference, contributing to personalization features, memory retrieval, and low‑latency serving in a fast‑paced startup environment based in San Francisco.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer - Scalable Training & Inference
ML Systems Engineer - Scalable Training & Inference

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Equity
Benefits
Commuter stipend
Senior ML Systems Engineer — Inference & Scale
Senior ML Systems Engineer — Inference & Scale

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
ML Systems & Performance Engineer
ML Systems & Performance Engineer

Engram • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive cash compensation
Startup equity
LLM Systems Engineer & Research
LLM Systems Engineer & Research

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Comprehensive health coverage
Dental and vision coverage
Retirement benefits
+3
Staff AI Systems Engineer — Inference & RL
Staff AI Systems Engineer — Inference & RL

Together • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Senior ML Inference Engineer: Production Systems
Senior ML Inference Engineer: Production Systems

MakerMaker • San Francisco (CA)

On-site
USD 180,000 - 240,000
Post-Training ML Systems Engineer for Next-Gen Agents
Post-Training ML Systems Engineer for Next-Gen Agents

Scale AI, Inc. • New York (NY)

On-site
USD 198,000 - 331,000
Health benefits
Equity
Learning and development stipend
+1
ML Systems Engineer: Scalable Training & Realtime Inference
ML Systems Engineer: Scalable Training & Realtime Inference

Jobzhr • New York (NY)

On-site
USD 180,000 - 280,000
Head of ML Systems & Inference
Head of ML Systems & Inference

Doist • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1