Member of Technical Staff (Evals & Post-Training)

Ambral Labs

New York (NY)

On-site

USD 165,000 - 325,000

Full time

8 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity and ownership
Equinox membership
Free meals, coffee, and snacks
Health insurance
Unlimited PTO

Job summary

Ambral Labs, based in New York, is seeking an engineer to build a replayable enterprise environment engine and related evaluation systems. You’ll work with the CTO to deploy into real workflows, observe model reasoning, and grade performance against real outcomes.

The role emphasizes ownership over research direction and production systems, with opportunities to work on reinforcement learning, agent harnesses, and large-scale datasets.

Qualifications

  • 1-7 years of experience building production software or machine-learning systems.
  • Bonus points for reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or closely related systems
  • Understand how environment design, reward design, context, tooling, and policy behavior interact
  • Turn fuzzy business objectives into tasks and signals that can be evaluated reliably
  • Diagnose whether a model’s limitations come from the model itself, its context, its tools, its harness, or its training
  • Move between research questions and production implementation without treating them as separate jobs
  • Write strong software and can build systems that process large, messy datasets at scale
  • Care about reproducibility, observability, and understanding why a model behaves the way it does
  • You’re looking to do the best work of your life and build something you’ll be proud of for decades

Responsibilities

  • Build an environment factory that converts recorded enterprise data and task definitions into runnable environments
  • Design graders that turn ambiguous business objectives into verifiable rewards
  • Develop methods for mining useful tasks, trajectories, and evaluation cases from historical workflows
  • Create eval sets that are representative, reproducible, and resistant to overfitting
  • Find the right combinations of models, tools, context, and policies to maximize performance while reducing inference cost
  • Train and evaluate agents that operate over long horizons, incomplete information, and large tool spaces
  • Build replay and observability systems that make agent behavior explainable and measurable
  • Scale from individual environments to thousands of concurrent training and evaluation runs

Skills

Production software
ML systems
Reinforcement learning
Evaluation infrastructure
Agent harnesses
LLM post-training

Job description

What we do

Ambral Labs helps enterprises own the intelligence behind their most important workflows. Every company has years of historical evidence showing how work gets done: the context people had, the decisions they made, the actions they took, and the outcomes that followed. Today, most of that history is inert. It isn’t structured in a way that companies can use to evaluate models and improve agent behavior. Ambral turns this history into replayable environments and eval sets grounded in real workflows and observed outcomes. We use those environments to improve model performance through reinforcement learning and other post-training techniques, alongside context engineering, harness design, and agent engineering. The result is better, more cost-efficient AI for each enterprise’s specific work, powered by open-weight models that the company owns and controls. This allows each company to retain ownership of its core intelligence instead of outsourcing it to a model provider. We graduated from YC S2025, raised millions in funding, and are already deployed within multi-billion dollar enterprises. Now we're growing the founding team

What You’ll Do

At the center of Ambral Labs is a replayable environment engine for the enterprise. The system reconstructs a company’s world as it existed at a particular moment in the past then exposes that state through the same tools an agent would use in production. This lets us place new policies and agent configurations inside real historical environments, observe how they reason and act, and grade their performance against real outcomes.

  • Building an environment factory that converts recorded enterprise data and task definitions into runnable environments
  • Designing graders that turn ambiguous business objectives into verifiable rewards
  • Developing methods for mining useful tasks, trajectories, and evaluation cases from historical workflows
  • Creating eval sets that are representative, reproducible, and resistant to overfitting
  • Finding the right combinations of models, tools, context, and policies to maximize performance while reducing inference cost
  • Training and evaluating agents that operate over long horizons, incomplete information, and large tool spaces
  • Building replay and observability systems that make agent behavior explainable and measurable
  • Scaling from individual environments to thousands of concurrent training and evaluation runs

These problems are wide open. You’ll have significant ownership over both the research direction and the production systems that make it real. You’ll work directly with the CTO, deploy into real enterprise workflows, and see your research tested against consequential problems and observable outcomes.

Who you are
  • You have 1-7 years of experience building production software or machine-learning systems (we're hiring at multiple levels for this role).
  • Bonus points for working on reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or closely related systems
  • You understand how environment design, reward design, context, tooling, and policy behavior interact
  • You’re comfortable turning fuzzy business objectives into tasks and signals that can be evaluated reliably
  • You can diagnose whether a model’s limitations come from the model itself, its context, its tools, its harness, or its training
  • You can move between research questions and production implementation without treating them as separate jobs
  • You write strong software and can build systems that process large, messy datasets at scale
  • You care about reproducibility, observability, and understanding why a model behaves the way it does
  • You’re looking to do the best work of your life and build something you’ll be proud of for decades

We care much more about what you’ve built and how you think than credentials or conventional career paths.

Benefits
  • Significant equity and ownership
  • Equinox membership
  • Free meals, coffee, and snacks
  • Health insurance
  • Unlimited PTO

Compensation Range: $165K - $325K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff (Evals & Post-Training)
Member of Technical Staff (Evals & Post-Training)

Ambral • New York (NY)

On-site
USD 120,000 - 180,000
Significant equity and ownership
Equinox membership
Free meals, coffee, and snacks
+2
Member of Technical Staff (New Grad)
Member of Technical Staff (New Grad)

Ambral (YC S25) • United States

On-site
USD 140,000 - 175,000
Equity ownership
Equinox membership
Free meals, coffee, and snacks
+2
Member of Technical Staff (New Grad)
Member of Technical Staff (New Grad)

Ambral (YC S25) • New York (NY)

On-site
USD 140,000 - 175,000
Equity ownership
Equinox membership
Free meals
+2
Head of Research
Head of Research

Ambral Labs • New York (NY)

On-site
USD 250,000 - 400,000
Equity and ownership
Equinox membership
Free meals, coffee, and snacks
+2
Member of Technical Staff (New Grad)
Member of Technical Staff (New Grad)

Ambral (YC S25) • San Francisco (CA)

On-site
USD 150,000 - 210,000
Significant equity and ownership
Equinox membership
Free meals, coffee, and snacks
+2
Head of Research
Head of Research

Ambral • New York (NY)

On-site
USD 180,000 - 280,000
Significant equity and ownership
Equinox membership
Free meals, coffee, and snacks
+2
Head of Research
Head of Research

Ambral Labs • San Francisco (CA)

On-site
USD 250,000 - 380,000
Equity
Equinox membership
Free meals
+2
Member of Technical Staff (New Grad)
Member of Technical Staff (New Grad)

Ambral • New York (NY)

On-site
USD 120,000 - 190,000
Significant equity and ownership
Equinox membership
Free meals, coffee, and snacks
+2
Head of Research (AI/ML)
Head of Research (AI/ML)

Ambral Labs • New York (NY)

Hybrid
USD 180,000 - 280,000
Head of Research (AI/ML)
Head of Research (AI/ML)

Ambral Labs • San Francisco (CA)

On-site
USD 180,000 - 290,000