Founding RL & Evals Engineer - Shape Long-Horizon AI

Runtime Labs, Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 200,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity stake
In-person collaboration

Job summary

Runtime Labs, Inc. is seeking a Founding Engineer to own RL and Evals for Óra, shaping how model behavior is measured across planning, memory, tool use, and long-horizon interaction.

You will build datasets, harnesses, and regression suites to drive product decisions, model selection, and post-training options when warranted. You will partner with cross-functional teams to define metrics, sampling, contamination control, and reproducible comparisons, ensuring evaluation signals guide engineering

Qualifications

  • Experience building evaluation systems for ML models or agents.
  • Strong background in RL, evaluation metrics, and regression testing.
  • Ability to design reproducible evaluation datasets and harnesses.

Responsibilities

  • Own the eval roadmap for RL and long-horizon model behavior.
  • Build datasets, annnotation schemas, and evaluation harnesses.
  • Drive product decisions through evaluation signals and experiments.

Skills

Evaluation systems
Reinforcement learning
Dataset design

Job description

Runtime Labs, Inc. is seeking a Founding Engineer to own RL and Evals for Óra, shaping how model behavior is measured across planning, memory, tool use, and long-horizon interaction.

You will build datasets, harnesses, and regression suites to drive product decisions, model selection, and post-training options when warranted. You will partner with cross-functional teams to define metrics, sampling, contamination control, and reproducible comparisons, ensuring evaluation signals guide engineering

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Founding Engineer, RL & Evals
Founding Engineer, RL & Evals

Runtime Labs, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Equity stake
In-person collaboration
RL Environment Engineer: Shape Frontier Model Learning
RL Environment Engineer: Shape Frontier Model Learning

Acceler8 Talent • San Francisco (CA)

On-site
USD 225,000 - 275,000
Equity
Senior Data & RL Engineer — Coding Agents Lead
Senior Data & RL Engineer — Coding Agents Lead

turing • Palo Alto (CA)

On-site
USD 250,000 - 350,000
Equity
Office in SF/Palo Alto/Seattle
Applied AI Engineer - Deploy, Debug & Build RL Evals
Applied AI Engineer - Deploy, Debug & Build RL Evals

HUD • San Francisco (CA)

On-site
USD 180,000 - 260,000
Medical, dental, vision benefits
Meals in office
Holiday break: Christmas Eve to New Y
+4
RL Research Scientist: Reasoning & Autoformalization
RL Research Scientist: Reasoning & Autoformalization

Pramaana Labs • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Senior Data & RL Environments Architect for AI Labs
Senior Data & RL Environments Architect for AI Labs

turing • San Francisco (CA)

On-site
USD 180,000 - 240,000
Founding Strategic Projects Lead for AI RL Labs
Founding Strategic Projects Lead for AI RL Labs

Morpheus Talent Solutions • New York (NY)

Hybrid
USD 200,000 - 250,000
Research Engineer, RL Environments and Infrastructure
Research Engineer, RL Environments and Infrastructure

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
RL Sandbox & Infra Engineer - Scalable Run-Time Platform
RL Sandbox & Infra Engineer - Scalable Run-Time Platform

Commergence • Fremont (CA)

Hybrid
USD 120,000 - 180,000
RL Environment Engineer: Shape Frontier AI
RL Environment Engineer: Shape Frontier AI

AI Talent Now • San Francisco (CA)

On-site
USD 150,000 - 250,000