ML Evaluation Engineer: RL Environments & Data Design

Wintermeyer Ventures

San Francisco (CA)

On-site

USD 260,000 - 290,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Wintermeyer Ventures is seeking an applied AI researcher to join our team in San Francisco. You will build Python-based software to support reinforcement learning environments and model evaluation, helping ML systems perform more reliably in real-world domains.

You will design data collection strategies and evaluation rubrics, create data slices to surface failure modes in finance, code, and enterprise workflows, and convert nuanced evaluation needs into structured data and assessment methods

Qualifications

  • 1–4 years of hands-on experience designing data collection strategies and evaluation rubrics for machine learning models.
  • Experience using Python for software development.
  • Background identifying model failure modes through evaluation across multiple domains, such as finance, code, or enterprise workflows.

Responsibilities

  • Develop Python software that supports reinforcement learning environment work and model evaluation.
  • Design data collection strategies and evaluation rubrics for machine learning models.
  • Create data slices that help surface model failure modes across finance, code, and enterprise workflows.
  • Help turn nuanced evaluation needs into structured data and assessment approaches that can be used to improve model performance.

Skills

Python
Data design
Reinforcement learning
Model evaluation

Job description

Wintermeyer Ventures is seeking an applied AI researcher to join our team in San Francisco. You will build Python-based software to support reinforcement learning environments and model evaluation, helping ML systems perform more reliably in real-world domains.

You will design data collection strategies and evaluation rubrics, create data slices to surface failure modes in finance, code, and enterprise workflows, and convert nuanced evaluation needs into structured data and assessment methods

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, RL Environments
Software Engineer, RL Environments

Wintermeyer Ventures • San Francisco (CA)

On-site
USD 260,000 - 290,000
AI Evaluation Engineer: RL Environments & Agents
AI Evaluation Engineer: RL Environments & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Research Engineer, RL Environments and Infrastructure
Research Engineer, RL Environments and Infrastructure

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
ML Research Engineer: Fine-Tuning & Evaluation Systems
ML Research Engineer: Fine-Tuning & Evaluation Systems

HonestAI • San Francisco (CA)

On-site
USD 140,000 - 190,000
Research Engineer – RL Infrastructure & Agent Environments
Research Engineer – RL Infrastructure & Agent Environments

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 150,000 - 190,000
Senior Software Engineer - RL Environments (Remote)
Senior Software Engineer - RL Environments (Remote)

YO AI Labs • Boston (MA)

Remote
USD 83,000 - 138,000
Remote work
Senior ML Engineer — Remote RL Environments
Senior ML Engineer — Remote RL Environments

YO AI Labs • San Francisco (CA)

Remote
USD 55,000 - 124,000
RL Environment Engineer: Shape Frontier AI
RL Environment Engineer: Shape Frontier AI

AI Talent Now • San Francisco (CA)

On-site
USD 150,000 - 250,000
ML Systems Engineer — RL Training & Finetuning
ML Systems Engineer — RL Training & Finetuning

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+1
Junior AI RL Environment Engineer — SF
Junior AI RL Environment Engineer — SF

Simplify • San Francisco (CA)

On-site
USD 255,000 - 345,000