SWE (RL Environments) "Reinforcement Learning"

AI Talent Now

San Francisco (CA)

On-site

USD 150,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AI Talent Now in San Francisco is seeking a skilled RL Environment Engineer to design impactful datasets that influence the learning of frontier AI models. You will collaborate with elite research teams at leading AI labs to refine reinforcement learning environments and create comprehensive benchmarks.

The ideal candidate will have a CS background from a top school and experience in reinforcement learning, with a strong emphasis on Python and Typescript. The position offers a competitive salary and bonus potential.

Qualifications

  • Recent graduates from top schools focused on excellence and depth.
  • First-author publications at top venues such as NeurIPS or ICML are highly desirable.
  • Experience in supervised fine-tuning (SFT) or reinforcement learning (RL).

Responsibilities

  • Design and develop datasets that shape frontier model learning.
  • Collaborate with research teams to build and refine RL environments.
  • Implement high-quality benchmarks and evaluate model performance.

Skills

Reinforcement learning environments
Python
Typescript

Education

CS degree from a top-30 school (US/CA/EUR)

Job description

San Francisco, United States | Posted on 06/15/2026

AI Talent Now, LLC is a powerhouse in direct‑hire talent acquisition, connecting exceptional talent with industry‑leading organizations across IT, Engineering, Financial Services & Fintech, and Manufacturing & Robotics. Headquartered in Atlanta, Georgia, we serve clients and candidates nationwide.

Job Description
AI Talent Now Job #ZR 72
About us

This dynamic company is helping push the frontier of LLMs and AI Agents through novel datasets and experimentation. We build the most complex infrastructure that powers frontier data creation for agentic and hard‑reasoning workflows. Working with all five leading AI labs, we are becoming the go‑to partner for data infrastructure for YC companies. Our sharp hockey‑stick growth and talent density come from a founding team with backgrounds in top IB and quant firms.

We build the training data and evaluation infrastructure that frontier AI labs use to improve their models. Working with the world's leading labs, we design high‑signal datasets and run rigorous evaluations that go beyond static benchmarks. At a small, early post‑Series A team, individual contributors directly influence how the next generation of models learn and improve.

As an RL Environment Engineer you will design datasets that directly influence how frontier models learn and work hands‑on with research teams at top AI labs.

Responsibilities
  • Design and develop datasets that shape frontier model learning.
  • Collaborate with research teams at leading AI labs to build and refine reinforcement learning environments.
  • Implement high‑quality benchmarks and evaluate model performance beyond static tests.
Qualifications
  • Recent graduates from top schools focused on excellence and depth; extensive track record not required.
  • First‑author publications at top venues such as NeurIPS or ICML are highly desirable.
  • Open to profiles from data companies with benchmarking experience.
  • Ideal candidates have created iconic benchmarks and possess experience in supervised fine‑tuning (SFT) or reinforcement learning (RL).
Experience
  • 1–6 years of experience as a software engineer.
  • Explicit experience building reinforcement learning environments.
  • Signals of excellence in software engineering (e.g., work at an RL company, top VC‑backed startup, strong side projects, quant or hedge‑fund background, or founding/early‑stage startup engineer).
  • CS degree from a top‑30 school (US/CA/EUR only).
  • Strong full‑stack skills with depth in Python, Typescript, and other backend languages.
  • Developed quantitative frameworks for measuring dataset quality/diversity.
  • Bias for action and execution; willing to tackle difficult and tedious work.
Compensation and Logistics
  • Base salary: $150k–$250k, with significant bonus potential based on performance.
  • Bonuses are uncapped and can substantially increase total compensation.
  • Visa sponsorships are possible (experience in H1B and other types).
Work Arrangement
  • Flexible working hours; most team members in the office from noon to midnight.
  • Work environment encourages results over strict hours.
EEO Statement

We are an equal opportunity employer and welcome applications from all qualified individuals regardless of race, color, religion, sex, gender identity, sexual orientation, national origin, disability, or veteran status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Environment Software Engineer
RL Environment Software Engineer

talentpluto • San Francisco (CA)

Hybrid
USD 180,000 - 220,000
Reinforcement Learning Environment Engineer
Reinforcement Learning Environment Engineer

Open Data Science • San Francisco (CA)

Remote
USD 100,000 - 150,000
RL Environment Engineer: Shape Frontier AI
RL Environment Engineer: Shape Frontier AI

AI Talent Now • San Francisco (CA)

On-site
USD 150,000 - 250,000
Research Engineer, Reinforcement Learning
Research Engineer, Reinforcement Learning

Techire Ai • San Francisco (CA)

On-site
USD 250,000 - 300,000
AfterQuery — Software Engineer, RL Environments
AfterQuery — Software Engineer, RL Environments

davidjoseph-co • San Francisco (CA)

On-site
USD 180,000 - 220,000
Competitive equity
Member of Technical Staff, Platform Engineering
Member of Technical Staff, Platform Engineering

David Joseph & Company • San Francisco (CA)

On-site
USD 200,000 - 250,000
Healthcare
Relocation support
401k with 4% match
+3
Research Engineer, Code RL (Reinforcement Learning) San Francisco, CA | New York City, NY
Research Engineer, Code RL (Reinforcement Learning) San Francisco, CA | New York City, NY

Anthropic • San Francisco (CA)

On-site
USD 500,000 - 850,000
Member of Technical Staff - Machine Learning Capabilities, New Graduates
Member of Technical Staff - Machine Learning Capabilities, New Graduates

Preference Model • San Francisco (CA)

On-site
USD 100,000 - 130,000
Competitive cash and equity compensation
Health, vision, dental benefits
401K match
+2
Member of Technical Staff - Machine Learning Capabilities, New Graduates
Member of Technical Staff - Machine Learning Capabilities, New Graduates

Preference Model • Seattle (WA)

On-site
USD 165,000 - 200,000
Competitive cash and equity compensation
Health, vision, dental benefits
401K match
+3
Research Engineer, Infrastructure, RL Systems
Research Engineer, Infrastructure, RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1