MACHINE LEARNING ENGINEER (GENERAL)

MakerMaker

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

MakerMaker in San Francisco seeks a senior ML engineer to own end-to-end ML systems, from data pipelines to deployment. You will translate research code into reliable infrastructure, ship at scale, and help the team move fast with trustworthy results.

You will collaborate daily with researchers, design observability, implement testing and runbooks, and set engineering standards so our experiments translate into repeatable, scalable production workloads.

Qualifications

  • 6+ years building production-grade ML systems.
  • Experience across data, training, evaluation, deployment, monitoring.
  • Strong distributed systems background.
  • Fluent in Python and ML frameworks (PyTorch or JAX).
  • Experience with experimentation infrastructure (Ray, Slurm, Kubernetes).

Responsibilities

  • Build and maintain the training, evaluation, and deployment pipelines that our research runs on.
  • Take research code from prototype to production: refactor, harden, instrument, test.
  • Design observability into our ML systems (metrics, logs, traces, eval dashboards).
  • Own data pipelines for training and evaluation: ingest, dedup, version, validate.
  • Work closely with researchers to understand needs and bottlenecks.
  • Set engineering standards across our ML stack (testing, reviews, runbooks).
  • Contribute to architectural decisions shaping research and production.

Skills

Python
Distributed Systems
ML Engineering
Research Collaboration
Observability

Tools

PyTorch
JAX
Ray
Slurm
Kubernetes

Job description

ABOUT THE COMPANY

We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site



ABOUT THE ROLE

You'll build and maintain the ML systems and pipelines that our research runs on top of data pipelines, training infrastructure, evaluation tooling, deployment, observability. The work bridges research and production, and you'll be the person who makes "we ran an experiment" actually mean "we ran it correctly, at scale, with results we trust."


This is a senior ML engineering role. You'll own systems end-to-end. You'll work with researchers daily and translate research code into infrastructure that the team can rely on. You'll move fast and you'll be measured on whether your systems make the team faster.



WHAT YOU'LL DO


  • Build and maintain the training, evaluation, and deployment pipelines that our research runs on

  • Take research code from prototype to production: refactor, harden, instrument, test

  • Design observability into our ML systems (metrics, logs, traces, eval dashboards) so failures surface fast

  • Own data pipelines for training and evaluation: ingest, dedup, version, validate

  • Work closely with researchers to understand what they need, what's slow, and what's brittle

  • Set engineering standards across our ML stack (testing, reviews, runbooks) so the team scales

  • Contribute to architectural decisions that shape how research and production interacts



WHAT WE'RE LOOKING FOR:


  • Senior ML engineer with 6+ years building production-grade ML systems

  • Track record across the full lifecycle: data, training, evaluation, deployment, monitoring

  • Strong distributed systems experience; you've shipped systems that have to be on

  • Fluent Python, fluent with at least one of (PyTorch, JAX); comfortable at the systems-level when needed

  • Comfortable with experimentation infrastructure (Ray, Slurm, Kubernetes, or similar)

  • Bias toward shipping; you prefer working code over working diagrams

  • Strong written communication



NICE TO HAVE:


  • Experience building experimentation platforms or research infrastructure at a frontier ML lab

  • Background in distributed training systems

  • Open-source contributions to ML infrastructure

  • History of working effectively with small senior teams



THIS ROLE IS PROBABLY NOT FOR YOU IF:


  • You want to do research with engineering as a side activity: this is engineering as the main thing

  • Cross-functional work with researchers (translation, scoping, education) doesn't appeal

  • Long-running ownership of running systems isn't appealing: this role has it

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MACHINE LEARNING ENGINEER (GENERAL)
MACHINE LEARNING ENGINEER (GENERAL)

MakerMaker.AI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 270,000
Lead Engineer, Machine Learning
Lead Engineer, Machine Learning

Salt Digital Recruitment • United States

On-site
USD 180,000 - 260,000
Machine Learning Engineer (Mid-Level)
Machine Learning Engineer (Mid-Level)

Clera • San Francisco (CA)

On-site
USD 120,000 - 150,000
ML Research Engineer
ML Research Engineer

Nodi • New York (NY)

On-site
USD 235,000 - 295,000
ML Research Engineer
ML Research Engineer

Radical AI • New York (NY)

On-site
USD 140,000 - 210,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

ExaCare AI • New York (NY)

On-site
USD 100,000 - 140,000
Flexible PTO
Medical, dental, and vision coverage
Company off-sites
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Regenerativeaitool • Boston (MA)

Hybrid
USD 120,000 - 150,000
Comprehensive health, dental, and vision insurance
Flexible PTO and remote-friendly work arrangements
Annual learning and development budget ($5,000)
+2
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Ultra • New York (NY)

Hybrid
USD 180,000 - 240,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Sierracorp • San Francisco (CA)

On-site
USD 150,000 - 200,000
Engineering Manager, ML
Engineering Manager, ML

cursor • San Francisco (CA), New York (NY)

On-site
USD 180,000 - 260,000