AI Evaluation Engineer: RL Environments & Agents

MaxIT Consulting - Max Corporate Group

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

MaxIT Consulting - Max Corporate Group seeks an AI Evaluation Engineer to design and build environments, evaluation systems, and infrastructure for training and assessing long-horizon enterprise AI agents. The role focuses on creating realistic agent environments and robust evaluation pipelines in a San Francisco setting.

The candidate will work across languages and stacks, shipping production-quality infrastructure and addressing reward design and trajectory analysis.

Qualifications

  • Hands-on experience with AI environments and evaluation systems.
  • Experience with RL infrastructure or agent-training systems.
  • Strong software engineering fundamentals.
  • Ability to build and ship production-grade infrastructure.
  • Knowledge of evaluation methods, reward design, and graders.

Responsibilities

  • Design evaluation environments for long-horizon enterprise agent workflows.
  • Define tasks, state, tools, graders, and reward signals for evaluation.
  • Build high-fidelity representations of complex enterprise software environments.
  • Develop infrastructure for rollouts, orchestration, and grader pipelines.
  • Measure correctness and efficiency across multi-step agent behavior.
  • Investigate evaluation failures and reward-quality issues.
  • Build production-quality systems beyond notebook research.

Skills

AI environments
Reinforcement learning infra
Software engineering
Infrastructure shipping
Evaluation methodology
Cross-language work

Job description

MaxIT Consulting - Max Corporate Group seeks an AI Evaluation Engineer to design and build environments, evaluation systems, and infrastructure for training and assessing long-horizon enterprise AI agents. The role focuses on creating realistic agent environments and robust evaluation pipelines in a San Francisco setting.

The candidate will work across languages and stacks, shipping production-quality infrastructure and addressing reward design and trajectory analysis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer – Reinforcement Learning & Agents
AI Evaluation Engineer – Reinforcement Learning & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Research Engineer, RL Environments and Infrastructure
Research Engineer, RL Environments and Infrastructure

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Remote AI Agent Evaluation Engineer
Remote AI Agent Evaluation Engineer

YO AI Labs • Los Angeles (CA)

Remote
USD 83,000 - 138,000
Remote Senior Software Engineer: AI Agent Evaluation
Remote Senior Software Engineer: AI Agent Evaluation

YO AI Labs • Washington

Remote
USD 60,000 - 90,000
RL Systems Engineer: Environments & High-Throughput Pipelines
RL Systems Engineer: Environments & High-Throughput Pipelines

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Remote AI Agent Evaluation Engineer
Remote AI Agent Evaluation Engineer

YO AI Labs • New York (NY)

Remote
USD 46,000 - 110,000
Remote Senior AI Agent Evaluation Engineer
Remote Senior AI Agent Evaluation Engineer

YO AI Labs • San Francisco (CA)

Remote
USD 83,000 - 124,000
Senior Software Engineer, AI Training & RL Environments
Senior Software Engineer, AI Training & RL Environments

YO AI Labs • North Carolina

Remote
USD 83,000 - 124,000
Senior AI Engineer - RL Environments for Codebase (Remote)
Senior AI Engineer - RL Environments for Codebase (Remote)

YO AI Labs • California (MO)

Remote
USD 83,000 - 124,000
Agent Systems AI Engineer - Production Deployments (SF)
Agent Systems AI Engineer - Production Deployments (SF)

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 170,000 - 250,000