Founding Engineer, RL & Evals

Runtime Labs, Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 200,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity stake
In-person collaboration

Job summary

Runtime Labs, Inc. is seeking a Founding Engineer to own RL and Evals for Óra, shaping how model behavior is measured across planning, memory, tool use, and long-horizon interaction.

You will build datasets, harnesses, and regression suites to drive product decisions, model selection, and post-training options when warranted. You will partner with cross-functional teams to define metrics, sampling, contamination control, and reproducible comparisons, ensuring evaluation signals guide engineering

Qualifications

  • Experience building evaluation systems for ML models or agents.
  • Strong background in RL, evaluation metrics, and regression testing.
  • Ability to design reproducible evaluation datasets and harnesses.

Responsibilities

  • Own the eval roadmap for RL and long-horizon model behavior.
  • Build datasets, annnotation schemas, and evaluation harnesses.
  • Drive product decisions through evaluation signals and experiments.

Skills

Evaluation systems
Reinforcement learning
Dataset design

Job description

# Founding Engineer, RL & EvalsRuntime Labs builds Óra at ora.app.As a Founding Engineer focused on RL & Evals, you will own how Runtime Labs measures model behavior in Óra: planning, scheduling, recommendations, memory use, tool use, and long-horizon interaction. You will define what good looks like, build the datasets and harnesses that score it, and close the loop so evaluation signal drives product decisions, model selection, and—when the data justifies it—post-training and reinforcement learning.This is not generic chatbot evaluation. The domain is agentic behavior over time: whether a model uses the right context, makes good plans, chooses appropriate actions, recovers from mistakes, and improves as interaction data compounds.## About the Role* Own the eval roadmap for planning, scheduling, recommendations, memory, tool use, and long-horizon model behavior in Óra* Define task- and system-level metrics that separate useful behavior from superficially plausible outputs—including temporal reasoning, retrieval quality, action selection, constraint satisfaction, and recovery from failure* Build and version evaluation datasets from synthetic scenarios, curated examples, production failures, and longitudinal interaction; turn real failures into durable regression cases* Design dataset schemas that preserve context, expected behavior, provenance, annotations, and outcomes—with clear conventions for sampling, contamination control, and reproducible comparisons* Build developer-friendly harnesses for writing, running, inspecting, and comparing evals; integrate them into model experimentation and everyday product engineering* Support offline replay and controlled comparison across models, prompts, retrieval strategies, and agent policies; make failures traceable from interaction through retrieval, reasoning, tool calls, and final action* Establish release criteria so model and product changes are measured before they reach users* Close the loop from eval signal to shipping decision: prioritize regressions, guide prompting/retrieval/architecture choices, and develop preference or reward signals from evaluation, user behavior, and longitudinal outcomes when warranted* Explore post-training and reinforcement-learning approaches when evaluation quality and interaction data justify them* Partner with Product, Infrastructure, and Data Infrastructure so production interaction becomes structured learning signal## About the Person* Research-minded engineer who cares deeply about measurement, experimental rigor, and why model behavior changes* Strong ML foundation with experience evaluating, training, or improving language-model or agentic systems* Moves between research questions and production engineering: define the experiment, build the dataset, implement the harness, analyze failures, ship the change* Strong intuition for evaluation design and the failure modes of automated metrics, model-based graders, and human labels* Reasons precisely about noisy signals, baselines, distributions, regressions, and experimental validity* Strong software engineer who builds durable evaluation infrastructure—not one-off notebooks or manual review alone* Interested in reinforcement learning, preference optimization, post-training, agent evaluation, and learning from interaction* Excited by planning, temporal reasoning, recommendations, memory, and models that act over long-lived user context* Wants high ownership as an early long-term partner defining how Runtime Labs measures and improves intelligent behavior## Experience* Experience building evaluation systems for LLMs, agents, recommendations, search, ranking, or decision-making systems* Experience with reinforcement learning, preference optimization, reward modeling, or post-training* Experience building datasets, annotation systems, graders, replay infrastructure, or experimentation platforms* Familiarity with model tracing, tool-use evaluation, retrieval evaluation, and multi-step agent benchmarks* Experience turning production model failures into reproducible test cases and measurable improvements* Background in sequential decision-making, recommendation systems, temporal reasoning, or human-AI interaction## Location & CompensationLocation: San Francisco. We are building the founding team in person and expect to work closely together during the company's early formation.Salary: $150k–$200k depending on experience and role.Equity: meaningful early-stage ownership for long-term partners.Exceptional candidates may be considered outside this range.We encourage you to apply even if you do not meet every qualification.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Founding RL & Evals Engineer - Shape Long-Horizon AI
Founding RL & Evals Engineer - Shape Long-Horizon AI

Runtime Labs, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Equity stake
In-person collaboration
Founding Engineer, Data
Founding Engineer, Data

Runtime Labs, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Equity
Founding Engineer, Infrastructure
Founding Engineer, Infrastructure

Runtime Labs, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Equity
Research Engineer, Policy Evaluation
Research Engineer, Policy Evaluation

Bonfirevc • Palo Alto (CA)

On-site
USD 150,000 - 190,000
AI Research Scientist, Learning & Evaluation
AI Research Scientist, Learning & Evaluation

Socket.dev • Beverly Hills (CA)

On-site
USD 180,000 - 280,000
Daily team dinner provided in-office
Research Manager, Evaluation
Research Manager, Evaluation

Aaru Inc. • New York (NY), Northern (KY)

Hybrid
USD 280,000 - 425,000
Equity
Visa sponsorship
Relocation support
+1
Research Engineer, Policy Evaluation
Research Engineer, Policy Evaluation

Orbifold AI, Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 230,000
Founding GTM & Partnerships
Founding GTM & Partnerships

Runtime Labs, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Research Engineer, Physical AI (Robotics, World Models)
Research Engineer, Physical AI (Robotics, World Models)

Orbifold AI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research Engineer - Evals
Research Engineer - Evals

Pantera Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive cash compensation
Equity opportunities
Relocation support
+1