Get more replies from employers
Send a job-specific resume in minutes.
Ambral is building a replayable environment engine over real enterprise history. The system reconstructs a company’s context from past data, exposing it through production-like tools so new policies and agent configurations can be tested against real outcomes.
You will own the research and infrastructure to turn this into a scalable model-improvement system, tackling environment design, graders, and evaluation while collaborating closely with the CTO on deployments.
Ambral helps enterprises own the intelligence behind their most important workflows.
Every company has years of historical evidence showing how work gets done: the context people had, the decisions they made, the actions they took, and the outcomes that followed. Today, most of that history is inert. It isn’t structured in a way that companies can use to evaluate models and improve agent behavior.
Ambral turns this history into replayable environments and eval sets grounded in real workflows and observed outcomes. We use those environments to improve model performance through reinforcement learning and other post-training techniques, alongside context engineering, harness design, and agent engineering.
The result is better, more cost-efficient AI for each enterprise’s specific work, powered by open-weight models that the company owns and controls. This allows each company to retain ownership of its core workflow intelligence instead of outsourcing it to a model provider.
We graduated from Y Combinator Summer 2025, raised millions in funding, and are already deploying within multi-billion-dollar enterprises. Now we’re growing the founding team.
We’re building a replayable environment engine over real enterprise history.
The system reconstructs a company’s context as it existed at any past time, then exposes that state through the same tools an agent would use in production. This lets us place new policies and agent configurations inside real historical environments, observe how they reason and act, and grade their performance against real outcomes.
You’ll own the research and infrastructure required to turn this into a scalable model-improvement system. The core problems include:
These problems are wide open. You’ll have significant ownership over both the research direction and the production systems that make it real.
You’ll work directly with the CTO, deploy into real enterprise workflows, and see your research tested against consequential problems and observable outcomes.
We care much more about what you’ve built and how you think than credentials or conventional career paths.
Compensation Range: $215K - $330K