AI Evaluation Engineer – Reinforcement Learning & Agents

MaxIT Consulting - Max Corporate Group

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

MaxIT Consulting - Max Corporate Group seeks an AI Evaluation Engineer to design and build environments, evaluation systems, and infrastructure for training and assessing long-horizon enterprise AI agents. The role focuses on creating realistic agent environments and robust evaluation pipelines in a San Francisco setting.

The candidate will work across languages and stacks, shipping production-quality infrastructure and addressing reward design and trajectory analysis.

Qualifications

  • Hands-on experience with AI environments and evaluation systems.
  • Experience with RL infrastructure or agent-training systems.
  • Strong software engineering fundamentals.
  • Ability to build and ship production-grade infrastructure.
  • Knowledge of evaluation methods, reward design, and graders.

Responsibilities

  • Design evaluation environments for long-horizon enterprise agent workflows.
  • Define tasks, state, tools, graders, and reward signals for evaluation.
  • Build high-fidelity representations of complex enterprise software environments.
  • Develop infrastructure for rollouts, orchestration, and grader pipelines.
  • Measure correctness and efficiency across multi-step agent behavior.
  • Investigate evaluation failures and reward-quality issues.
  • Build production-quality systems beyond notebook research.

Skills

AI environments
Reinforcement learning infra
Software engineering
Infrastructure shipping
Evaluation methodology
Cross-language work

Job description

San Francisco, California | Primarily On-site

We are seeking an AI Evaluation Engineer – Reinforcement Learning & Agents to build the environments, evaluation systems, and supporting infrastructure used to train and assess long-horizon enterprise AI agents.

The Opportunity

You will work on the engineering and research problems behind realistic agent environments, post-training systems, and reliable evaluation of complex multi-step workflows.

Key Responsibilities
  • Design evaluation environments for long-horizon enterprise agent workflows.
  • Define tasks, state, tools, graders, and reward signals used to evaluate and improve agents.
  • Build high-fidelity representations of complex enterprise software environments.
  • Develop infrastructure for rollouts, orchestration, trajectory inspection, and grader pipelines.
  • Measure both correctness and efficiency across multi-step agent behavior.
  • Investigate evaluation failures, reward-quality issues, and agent behavior.
  • Build production-quality systems rather than notebook-only research prototypes.
Required Qualifications
  • Hands‑on experience with AI environments, evaluations, reinforcement learning infrastructure, or related agent‑training systems.
  • Strong software engineering fundamentals.
  • Demonstrated ability to build and ship technical infrastructure.
  • Understanding of evaluation methodology, reward design, graders, and agent trajectories.
  • Ability to work across languages and technology stacks based on system requirements.
Candidate Profile

A PhD is not required. Strong engineering and shipped environment or evaluation systems are more important than academic credentials or publication history.

Seniority

The opportunity is open to exceptional new graduates, early‑career engineers, and experienced senior candidates. Selection is based primarily on engineering strength and relevant technical work.

Work Arrangement

The role is anchored in San Francisco with a strong preference for in‑person collaboration. Limited flexibility may be considered case by case for exceptional candidates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer: RL Environments & Agents
AI Evaluation Engineer: RL Environments & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Forward Deployed AI Engineer – Agent Systems
Forward Deployed AI Engineer – Agent Systems

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 170,000 - 250,000
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
Research Engineer, RL Environments and Infrastructure
Research Engineer, RL Environments and Infrastructure

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
AI Engineer, Evals & Agent Quality
AI Engineer, Evals & Agent Quality

Town.com, Inc. • San Francisco (CA)

On-site
USD 250,000 - 300,000
Remote Senior Software Engineer: AI Agent Evaluation
Remote Senior Software Engineer: AI Agent Evaluation

YO AI Labs • Washington

Remote
USD 60,000 - 90,000
Agentic AI Engineer
Agentic AI Engineer

Oscar • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Agent Engineer San Francisco CA
Agent Engineer San Francisco CA

AHU Technologies Inc • Washington

On-site
USD 120,000 - 150,000
Agent Post-Training, Frontier Evals and Environments Research
Agent Post-Training, Frontier Evals and Environments Research

OpenAI • San Francisco (CA)

On-site
USD 380,000 - 500,000
Research Engineer — AI Alignment & Evaluation
Research Engineer — AI Alignment & Evaluation

W3 Sourcing • San Francisco (CA)

Hybrid
USD 140,000 - 210,000