ML ENGINEER (GENERAL)

MakerMaker

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MakerMaker in San Francisco is seeking a senior ML engineer to own end-to-end ML systems, from data pipelines to deployment. You’ll translate research code into robust infrastructure, ship features at scale, and ensure results are trustworthy.

You’ll collaborate with researchers daily, implement observability dashboards, raise the bar on testing and runbooks, and drive architectural decisions that keep our research and production aligned and fast.

Qualifications

  • Senior ML engineer with 6+ years building production-grade ML systems.
  • Track record across data, training, evaluation, deployment, and monitoring.
  • Strong distributed systems experience; shipped systems at scale.
  • Fluent in Python; comfortable with PyTorch or JAX; systems-level know-how.
  • Experience with experimentation infrastructure (Ray, Slurm, Kubernetes).
  • Bias toward shipping and strong written communication.

Responsibilities

  • Build and maintain the training, evaluation, and deployment pipelines that our research runs on.
  • Take research code from prototype to production: refactor, harden, instrument, test.
  • Design observability into ML systems (metrics, logs, traces, dashboards) so failures surface fast.
  • Own data pipelines for training and evaluation: ingest, dedup, version, validate.
  • Work closely with researchers to understand needs, bottlenecks, and brittleness.
  • Set engineering standards across the ML stack (testing, reviews, runbooks) to scale the team.
  • Contribute to architectural decisions that shape research and production interaction.

Skills

Distributed systems
Python
MLOps
Data pipelines
Strong written communication

Tools

PyTorch
JAX
Ray
Kubernetes
Slurm

Job description

ABOUT THE COMPANY

We’re building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We’re a small team based in San Francisco, on-site


ABOUT THE ROLE

You’ll build and maintain the ML systems and pipelines that our research runs on top of: data pipelines, training infrastructure, evaluation tooling, deployment, observability. The work bridges research and production, and you’ll be the person who makes \"we ran an experiment\" actually mean \"we ran it correctly, at scale, with results we trust.\"


This is a senior ML engineering role. You’ll own systems end-to-end. You’ll work with researchers daily and translate research code into infrastructure that the team can rely on. You’ll move fast and you’ll be measured on whether your systems make the team faster.


WHAT YOU’LL DO


  • Build and maintain the training, evaluation, and deployment pipelines that our research runs on


  • Take research code from prototype to production: refactor, harden, instrument, test


  • Design observability into our ML systems (metrics, logs, traces, eval dashboards) so failures surface fast


  • Own data pipelines for training and evaluation: ingest, dedup, version, validate


  • Work closely with researchers to understand what they need, what’s slow, and what’s brittle


  • Set engineering standards across our ML stack (testing, reviews, runbooks) so the team scales


  • Contribute to architectural decisions that shape how research and production interact



WHAT WE’RE LOOKING FOR


  • Senior ML engineer with 6+ years building production-grade ML systems


  • Track record across the full lifecycle: data, training, evaluation, deployment, monitoring


  • Strong distributed systems experience; you’ve shipped systems that have to be on


  • Fluent Python, fluent with at least one of (PyTorch, JAX); comfortable at the systems-level when needed


  • Comfortable with experimentation infrastructure (Ray, Slurm, Kubernetes, or similar)


  • Bias toward shipping; you prefer working code over working diagrams


  • Strong written communication



NICE TO HAVE


  • Experience building experimentation platforms or research infrastructure at a frontier ML lab


  • Background in distributed training systems


  • Open-source contributions to ML infrastructure


  • History of working effectively with small senior teams



THIS ROLE IS PROBABLY NOT FOR YOU IF


  • You want to do research with engineering as a side activity: this is engineering as the main thing


  • Cross-functional work with researchers (translation, scoping, education) doesn’t appeal


  • Long-running ownership of running systems isn’t appealing: this role has it


Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RESEARCH ENGINEER (GENERAL)
RESEARCH ENGINEER (GENERAL)

MakerMaker.AI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Senior ML Engineer
Senior ML Engineer

Next Ventures • New York (NY)

On-site
USD 130,000 - 160,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

ExaCare AI • New York (NY)

On-site
USD 100,000 - 140,000
Flexible PTO
Medical, dental, and vision coverage
Company off-sites
RESEARCHER (GENERAL)
RESEARCHER (GENERAL)

MakerMaker • San Francisco (CA)

On-site
USD 150,000 - 210,000
Engineering Manager, ML
Engineering Manager, ML

Cursor • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior ML Infra Engineer
Senior ML Infra Engineer

Maxinsights Corporation • Santa Clara (CA)

On-site
USD 180,000 - 240,000
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Sierracorp • San Francisco (CA)

On-site
USD 150,000 - 200,000
Backend Software Engineer (ML Infra)
Backend Software Engineer (ML Infra)

Rockstar • San Francisco (CA)

On-site
USD 100,000 - 130,000
Engineering Manager, ML
Engineering Manager, ML

Cursor • New York (NY)

On-site
USD 180,000 - 260,000