RL Environments Architect

Surge AI

United States

Remote

USD 150,000 - 230,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Surge AI is seeking an RL Environments Architect to design and govern simulated worlds for training agents, from simple task microcosms to multi-agent ecosystems.

You will define primitives, rewards, interfaces and telemetry, ensure data quality and reproducibility, and collaborate with researchers to translate real-world tasks into robust simulations.

Qualifications

  • Simulation & systems depth: building RL environments/simulators with determinism, performance, observability.
  • Data quality leadership: design reward functions, taxonomy, QA to keep signals aligned.
  • Builder's mindset: collaborate across research and engineering to ship pragmatic, testable environments.

Responsibilities

  • Architect a modular environment framework with clear APIs, curricula, and configurable reward/termination schemas.
  • Establish quality bars: coverage metrics, invariance checks, and trace audits for environment outputs and agent experience buffers.
  • Instrument telemetry for episode rollouts; mitigate reward hacking, mode collapse, and exploitable loopholes.
  • Partner with researchers to translate real-world tasks into robust simulations, including synthetic data generators and evaluation suites.

Job description

About Us

Our mission is to raise AGI with the richness of human intelligence — curious, witty, imaginative, and full of unexpected brilliance.

Surge was founded by engineers and researchers who dreamed of building the next generation AI. We're building a platform that powers the most powerful models in the world in partnership with companies like Anthropic, Google, Microsoft, and Meta.

At Surge, we believe the path to AGI isn't just about scaling compute—it's about embracing the unlimited ceiling of human intelligence and creativity in the data that shapes these systems. Our platform combines elite human expertise with cutting-edge tools for scalable oversight, from building rich RL environments to conducting rigorous evaluations that go beyond benchmarks. We've run a profitable business from day one without raising venture funding.

The Role

As an RL Environments Architect, you’ll design, instrument, and govern the simulated worlds where agents learn — from compact task microcosms to multi-agent, tool-using ecosystems. You’ll define the primitives, reward structures, interfaces, and telemetry that let us stress-test emerging capabilities while keeping training signals faithful, stable, and scalable.

Not only will you build environments, you’ll craft standards for data quality and reproducibility across large-scale agent gyms. This is a role for someone who sweats the details of simulation fidelity, thinks in terms of coverage and failure surfaces, and loves turning messy real-world phenomena into learnable curricula. Your work will form the backbone for safe, rapid progress in agentic systems.

What You'll Do
  • Architect a modular environment framework with clear APIs, curriculum scaffolds, and configurable reward/termination schemas

  • Establish quality bars: coverage metrics, invariance checks, and trace audits for environment outputs and agent experience buffers

  • Instrument rich telemetry for episode rollouts; mitigating reward hacking, mode collapse, and exploitable loopholes

  • Partner with researchers to translate real-world tasks into robust simulations, including synthetic data generators and evaluation suites

What We’re Looking for
  • Simulation & Systems Depth – Experience building RL environments or simulators (e.g., custom physics, multi-agent, tool APIs) with an eye for determinism, performance, and observability

  • Data Quality Leadership – Strong instincts for designing reward functions, scenario taxonomies, and QA pipelines that keep signals aligned and drift-free

  • Builder’s Mindset – Comfort collaborating across research and engineering to ship pragmatic, testable environments that evolve with model capabilities

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

RL Environments Engineer
RL Environments Engineer

Bespoke-Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision
401(k)
Daily onsite lunch
+3
RL Environments Engineer
RL Environments Engineer

Bespoke Labs Inc. • Mountain View (CA), Northern (KY)

On-site
USD 250,000 - 300,000
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
+2
Forward Deployed Engineer, RL Environments
Forward Deployed Engineer, RL Environments

WeHireYou • San Francisco (CA)

On-site
USD 140,000 - 190,000
Human Data Architect
Human Data Architect

Surge AI • United States

Remote
USD 120,000 - 210,000
Research Engineer, Coding Evaluation & Training Data
Research Engineer, Coding Evaluation & Training Data

Surge AI • United States

Remote
USD 140,000 - 210,000
Member of Technical Staff - RL Environments
Member of Technical Staff - RL Environments

Cohere • United States

Remote
USD 180,000 - 240,000
Research Engineer, RL Environments and Infrastructure
Research Engineer, RL Environments and Infrastructure

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
RL Environment Software Engineer
RL Environment Software Engineer

talentpluto • San Francisco (CA)

On-site
USD 180,000 - 220,000
Research Scientist
Research Scientist

Surge AI • United States

Remote
USD 130,000 - 210,000
AI Lab Tech Engineer
AI Lab Tech Engineer

Commergence • Colorado

On-site
USD 150,000 - 210,000