Staff AI/Machine Learning Engineer

Cacheflow

San Francisco (CA)

Hybrid

USD 180,000 - 260,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive salary and equity
Unlimited paid time off
401k plan with employer contribution
Medical, dental, and vision insurance
Generous parental leave policy
Remote-friendly work environment

Job summary

Tonic builds the data infrastructure behind modern AI. The models you build here are load-bearing, impacting whether agents are ready to ship and how realistic synthetic environments are.

You’ll work on synthesis and de-identification models with real enterprise data and high-stakes challenges that defy textbook answers. You’ll design scalable systems for synthetic environments, train and evaluate models at scale, and collaborate with frontier labs and enterprise ML teams to drive concrete model

Qualifications

  • 8+ years (or PhD with 3+ years) building production ML systems.
  • Experience with LLMs, agents, RL, NER, or information extraction.
  • Hands-on experience training and shipping models to production.
  • Fluency with modern training/eval stacks (PyTorch, distributed training).

Responsibilities

  • Design and build the systems that generate longitudinally coherent synthetic environments for agent training and evaluation.
  • Build synthesis models that generate realistic replacement values at large scale, preserving format and distribution.
  • Train and improve the NER models behind entity detection across free text and structured data.
  • Build evaluation infrastructure that grades agent outcomes on real tasks.
  • Fine-tune open-weight models on Tonic-generated data and turn results into direction.
  • Expand coverage into new domains, languages, and entity types.
  • Own model evaluation across precision/recall and outcome-level grading.
  • Optimize inference for large volumes of sensitive data inside customer environments.
  • Partner with frontier labs and enterprise ML teams to ship model improvements.
  • Set technical direction for a small, senior team.

Skills

LLMs
Agents
Reinforcement Learning
NER
Information extraction

Education

PhD inCS/ML

Tools

PyTorch
Distributed training
Benchmark frameworks

Job description

About Tonic

Tonic builds the data infrastructure behind modern AI. We generate the synthetic environments that agents are trained and tested in, and we de-identify real enterprise data so it can be used safely in training and evaluation. Eight years in, we work with frontier AI labs pushing the edge of what models can do, and with hundreds of enterprises including Fidelity, Comcast, eBay, and Vanguard, on the data problems that sit at the center of where AI is going next.


About The Role

The models you build here are load-bearing. The environments you generate decide whether an agent is ready to ship or only looked good in a demo. The synthesis and de-identification models you train decide whether a bank can safely put its data near a model at all. And the work spans real range: in one week you might build evaluation that separates the best models from the rest on real tasks, train a synthesis model where both fidelity and downstream utility have to hold, and improve entity detection on messy production data. Real enterprise data, real stakes, and problems that don’t have textbook answers yet.


What You'll Do


  • Design and build the systems that generate longitudinally coherent synthetic environments for agent training and evaluation, including persona modeling, task generators, and verifiable ground truth.


  • Build and maintain synthesis models that generate realistic replacement values at very large scale, preserving format, statistical distribution, and semantic consistency so de-identified data stays useful downstream.


  • Train and improve the NER models behind our entity detection, driving accuracy and recall across free text, structured fields, and mixed enterprise data at scale.


  • Build evaluation infrastructure that grades agent outcomes, not just traces, and produces real discrimination between frontier models on real tasks.


  • Fine-tune and evaluate open-weight models on Tonic-generated data, and turn benchmark results into product and research direction.


  • Expand coverage into new domains, languages, and entity types, and handle the long tail of formats and edge cases that real customer data throws off.


  • Own model evaluation across the board: precision and recall on detection, utility preservation on synthesis, and outcome-level grading for agents.


  • Optimize inference so models run efficiently on large volumes of sensitive data inside customer environments.


  • Partner directly with frontier labs and enterprise ML team to turn hard data problems into shipped model improvements.


  • Set technical direction for a small, senior team and raise the bar on rigor, reproducibility, and shipping.



What You'll Bring


  • 8+ years (or PhD with 3+ years) building production ML systems, with real depth in some combination of LLMs, agents, RL, NER, or information extraction.


  • Hands-on experience training and shipping models to production, and a pragmatic bar for quality: you know how to measure it, where it breaks, and when it's good enough to ship.


  • Experience with generative or synthesis models where output fidelity and downstream utility both matter, not just plausibility.


  • Strong software engineering fundamentals. You write code others build on.


  • Fluency with modern training and eval stacks (PyTorch, distributed training, standard agent and benchmark frameworks).


  • Comfort working with messy, sensitive, real-world data and the privacy constraints that come with it.


  • A track record of framing ambiguous problems and driving them to measurable, shipped results.


  • Bonus: synthetic data generation, data privacy or de-identification, or benchmark construction.



Benefits We Offer


  • Competitive salary and equity


  • Unlimited paid time off


  • 401k plan with employer contribution


  • Medical, dental, and vision insurance


  • Generous parental leave policy


  • Remote-friendly work environment


Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI & ML Engineer - Remote, Unlimited PTO
Senior AI & ML Engineer - Remote, Unlimited PTO

Tonic AI • United States

On-site
USD 180,000 - 240,000
Competitive salary and equity
Unlimited PTO
401k with employer contribution
+3
Senior AI & ML Engineer — Remote-friendly, Synthetic Environments
Senior AI & ML Engineer — Remote-friendly, Synthetic Environments

Cacheflow • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Competitive salary and equity
Unlimited paid time off
401k plan with employer contribution
+3
Design Engineer
Design Engineer

Applied Compute • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive compensation and equity
Generous health benefits
Unlimited PTO
+4
Founding AI Engineer, Agents & Evaluation
Founding AI Engineer, Agents & Evaluation

Zingage • New York (NY)

On-site
USD 190,000 - 240,000
Competitive base and meaningful equity
Equipment stipend
Luxury gym membership in NYC
+4
Member of Technical Staff (New Grad)
Member of Technical Staff (New Grad)

Ambral • New York (NY)

On-site
USD 120,000 - 190,000
Significant equity and ownership
Equinox membership
Free meals, coffee, and snacks
+2
Senior Backend Engineer
Senior Backend Engineer

In Tandem • Minnesota

On-site
USD 150,000 - 210,000
Medical premium coverage for employees
401k match
Paid parental leave
+3
Remote AI/NLP Engineer for Gen AI & Scalable ML
Remote AI/NLP Engineer for Gen AI & Scalable ML

Agent IQ • United States

Remote
USD 90,000 - 120,000
Agent IQ swag
Great teammates
Member of the Technical Staff - Systems ML Engineer
Member of the Technical Staff - Systems ML Engineer

Breakout Ventures • Cambridge (MA)

On-site
USD 180,000 - 270,000
Equity
Lunch subsidy
Health insurance
+1
AI Research Scientist, Learning & Evaluation
AI Research Scientist, Learning & Evaluation

Studyfetch • Beverly Hills (CA)

On-site
USD 150,000 - 210,000
Medical, Dental, Vision (100% employer
75% dependent coverage
401(k) with employer matching
+2
Member of Technical Staff (Evals & Post-Training)
Member of Technical Staff (Evals & Post-Training)

Ambral • New York (NY)

On-site
USD 120,000 - 180,000
Significant equity and ownership
Equinox membership
Free meals, coffee, and snacks
+2