RL Evaluation Engineer: Build Production Benchmarks

Invisible, Inc.

Northern, New York (KY, NY)

Hybrid

USD 150,000 - 210,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Bonuses
Equity

Job summary

Invisible Technologies is building a reinforcement learning capability within its research organization. As a Research Engineer, you will design evaluation methods, build scoring frameworks, and ship production systems that measure frontier models for enterprise clients.

You’ll influence methodology, not just implementation, and collaborate across ML, solutions, and research teams to drive impact. This role emphasizes turning ambiguous research questions into working systems with

Qualifications

  • Production-quality code written daily; this is a hard requirement.
  • Fluency in Python and comfort across the modern ML stack.
  • Experience building evaluation systems, RL environments, or training and inference infrastructure.
  • Hands-on experience with modern agentic flows.
  • Familiarity with reinforcement learning methods and frontier evaluation.

Responsibilities

  • Design benchmarks and RL environments to measure real model capability.
  • Originate evaluation methodology and translate into scoring frameworks and architectures.
  • Write and ship production code implementing designs.
  • Build and maintain data pipelines for evaluation runs.
  • Run analyses to verify results and ensure reproducibility.
  • Collaborate with Research Scientists, Solutions Architects, and ML engineers.

Skills

Python
ML Stack
Reinforcement Learning
Agentic Systems
Research to Production

Tools

Git
Docker
Kubernetes

Job description

Invisible Technologies is building a reinforcement learning capability within its research organization. As a Research Engineer, you will design evaluation methods, build scoring frameworks, and ship production systems that measure frontier models for enterprise clients.

You’ll influence methodology, not just implementation, and collaborate across ML, solutions, and research teams to drive impact. This role emphasizes turning ambiguous research questions into working systems with

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer — RL Evaluation & Benchmarks
Research Engineer — RL Evaluation & Benchmarks

Invisible Technologies Inc. • New York (NY)

Hybrid
USD 150,000 - 210,000
Bonuses and equity
Hybrid work environment
Evaluation Platform Engineer: Build Scalable ML Benchmarks
Evaluation Platform Engineer: Build Scalable ML Benchmarks

Thinking Machines Lab • San Francisco (CA)

On-site
USD 300,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
AI Evaluations Engineer — Benchmarking Frontiers
AI Evaluations Engineer — Benchmarking Frontiers

Meta • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Remote ML Engineer: Build Benchmarks & Production Pipelines
Remote ML Engineer: Build Benchmarks & Production Pipelines

raydar • Northern (KY)

Hybrid
USD 170,000 - 270,000
Competitive equity
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
Benchmarking Research Engineer: Frontier Model Evaluations
Benchmarking Research Engineer: Frontier Model Evaluations

Refresh AI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Research Engineer, Post-Training Evaluation & Infra
Research Engineer, Post-Training Evaluation & Infra

Ersilia • San Francisco (CA)

On-site
USD 150,000 - 350,000
Meaningful equity grants
Health, dental, and vision coverage
Staff Engineer - AI Evaluation & Metrics Platform
Staff Engineer - AI Evaluation & Metrics Platform

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 200,000
Research Scientist
Research Scientist

Anyone AI Inc. • Northern (KY)

On-site
USD 110,000 - 160,000
ML Research Engineer: Fine-Tuning & Evaluation Systems
ML Research Engineer: Fine-Tuning & Evaluation Systems

HonestAI • San Francisco (CA)

On-site
USD 140,000 - 190,000