MTS, Post-Training (Enterprise)

Bespoke-Labs

Mountain View (CA)

On-site

USD 300,000 - 350,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Visa sponsorship
Relocation support
Onsite daily lunch
Health, dental, and vision coverage
Equity
Performance-based bonus
Remote considered

Job summary

Bespoke Labs is seeking a delivery-focused machine learning practitioner in Mountain View, CA, to stand up our enterprise post-training capability. You will post-train, fine-tune, and align models for real-world business domains, ensuring reliable production performance and clear value delivery to enterprise customers.

You will own the evals, build bespoke benchmarks, curate data flywheels, and work directly with stakeholders to translate requirements into measurable metrics.

Qualifications

  • Experience shipping models to production and measuring success.
  • Proven ownership of end-to-end evals and benchmarks.
  • Deep understanding of regression risks when updating models.
  • Customer-facing collaboration with enterprise stakeholders.
  • Strong software and ML fundamentals for production deployments.
  • Owns full post-training lifecycle from data to deployment.

Responsibilities

  • Ship enterprise-grade models post-training and fine-tune for production.
  • Build bespoke evaluation suites and calibrate judges against experts.
  • Curate and filter high-impact datasets with strict quality controls.
  • Manage and mitigate regression risk across deployments.
  • Engage with product and enterprise customers to translate requirements into metrics.
  • Deploy with cost and privacy considerations for enterprise use.
  • Direct frontier tools and workflows to maximize output quality and speed.

Skills

Shipped models experience
End-to-end eval ownership
Regression risk understanding
Customer-facing / product empathy
Software & ML fundamentals
Ownership mindset

Job description

About Bespoke Labs

Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents.

Recently, we curated Open Thoughts, one of the best open reasoning datasets used by multiple frontier labs, trained SOTA specialized models such as Bespoke-MiniChart-7B and Bespoke-MiniCheck, and taught agents to do multi-turn tool-calling with reinforcement learning.

Bespoke is uniquely positioned to capture a large market share of data and RL environment curation.

About the Role

This is a delivery-heavy role focused on standing up our enterprise post-training capability. Demand is inbound, and our goal is to ship custom, high-performing models to 2–3 paying enterprise customers by year-end. Over the first 6–12 months, you will do the execution work required to solve real-world enterprise problems while helping build the scalable product underneath.

You will not be running abstract experiments or training models solely for benchmarks. You will sit directly at the intersection of enterprise demand and applied post-training—curating data, building rigorous eval suites, fine-tuning models, and proving their value to enterprise stakeholders. We will measure you on the production impact, robustness, and delivery of models shipped to real users, not on published papers.

The thing we care about most is whether you have done this before. If you have post-trained an LLM, shipped it to production users, managed regression risks, and owned the evals from end-to-end, we want to talk.

What You'll Do

  • Ship enterprise-grade models. Post-train, fine-tune, and align open-weight and proprietary base models for complex business domains, ensuring they perform reliably in production.

  • Build bespoke evaluation suites. Define what "quality" means for subjective domain-specific tasks, create custom benchmarks, and calibrate LLM judges against human domain experts.

  • Curate and filter high-impact datasets. Build and run production data flywheels combining real production traces, human labeling, and synthetic augmentation with strict filtering standards.

  • Manage and mitigate regression risk. Rigorously track downstream performance to ensure fine-tuning for new behaviors doesn't silently degrade core capabilities or reasoning.

  • Engage directly with stakeholders. Sit in front of product and enterprise customers to understand their requirements, translate vague domain preferences into technical eval metrics, and explain model behavior clearly.

  • Deploy for cost and privacy. Fine-tune open-weight architectures (e.g., Llama, Qwen, Mistral, DeepSeek) to hit strict enterprise latency, cost, and privacy targets.

  • Direct frontier tools and workflows. Leverage state-of-the-art post-training techniques, preference tuning, and data curation tooling to maximize output quality and delivery speed.

What We're Looking For

  • A record of shipped models. You have post-trained at least one LLM that was deployed to real users in production, and you can show how you measured its success.

  • End-to-end eval ownership. Demonstrated experience building benchmarks, creating eval datasets, and getting stakeholders to agree on clear metrics for complex or subjective tasks.

  • Deep understanding of regression risks. You can instinctively explain how training a model on new behaviors impacts existing capabilities and how to prevent it.

  • Customer-facing or product empathy. Experience collaborating directly with non-ML stakeholders, enterprise customers, or product managers to turn requirements into model behavior.

  • Strong software and ML fundamentals. Fluency in modern post-training frameworks, data processing pipelines, and code bases designed for production deployment.

  • Ownership mindset. You take complete responsibility for the full post-training lifecycle—from raw data to model deployment and failure analysis—without requiring close supervision.

You May Be a Good Fit If You Also

  • Have post-trained conversational or task-oriented assistants (e.g., support agents, multi-turn chat, tool-using agents)

  • Have built LLM judges or reward models and calibrated them against human raters

  • Have operated a production data flywheel: traces → labeling → synthetic augmentation → retrain

  • Have extensive hands-on experience with open-weight models (Llama, Qwen, Mistral, DeepSeek) for cost, latency, or privacy optimization

  • Come from forward-deployed engineer (FDE), founder, or early-stage startup backgrounds

What We Offer

  • Location: Mountain View, CA (Preferred) or San Francisco, CA (Onsite); Remote considered

  • Base Salary: $300,000 – $350,000 USD / year

  • Additional Comp: 25% performance-based bonus + equity

Benefits & Perks:

  • Health, dental, and vision coverage

  • 401(k)

  • Daily onsite lunch provided

  • Visa sponsorship and relocation support available

  • Direct impact on how the industry trains and evaluates agents

We value different backgrounds and paths into this work. If this role excites you but you do not check every box, apply anyway.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MTS, Post-Training (Enterprise)
MTS, Post-Training (Enterprise)

Bespoke Labs Inc. • Mountain View (CA), Northern (KY)

Hybrid
USD 300,000 - 350,000
Location Mountain View, CA (Onsite)
Base salary $300,000–$350,000 USD/year
25% performance-based bonus
+3
RL Environments Engineer
RL Environments Engineer

Bespoke-Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision
401(k)
Daily onsite lunch
+3
RL Environments Engineer
RL Environments Engineer

Bespoke Labs Inc. • Mountain View (CA), Northern (KY)

On-site
USD 250,000 - 300,000
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
+2
Research Engineer, Training and Environment Infrastructure
Research Engineer, Training and Environment Infrastructure

Ersilia • San Francisco (CA)

On-site
USD 150,000 - 350,000
Meaningful equity grants
Health, dental, and vision coverage
ML Engineer, Post Training
ML Engineer, Post Training

Varick Agents LTD. • San Francisco (CA)

On-site
USD 150,000 - 190,000
Equity
Flexible PTO
Free lunch & dinner in office
+2
Member of Technical Staff, Post-Training
Member of Technical Staff, Post-Training

Goaly • Menlo Park (CA), Northern (KY)

On-site
USD 150,000 - 230,000
Meals and office benefits
Visa sponsorship
Member of Technical Staff - Research Engineer, Post-training
Member of Technical Staff - Research Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive cash and equity (>90th pct
Ownership and autonomy
Lunch onsite
+4
Forward Deployed Engineer, Lead - LLM Post-training
Forward Deployed Engineer, Lead - LLM Post-training

Reflection AI • New York (NY)

On-site
USD 180,000 - 240,000
Stock options
Health insurance
Meals provided
+4
Machine Learning Researcher
Machine Learning Researcher

Multicoin • San Francisco (CA)

On-site
USD 250,000 - 350,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
ML Ops Engineer — Agentic AI Lab (Founding Team)
ML Ops Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Meaningful equity