AI Evaluation Engineer – Reinforcement Learning & Agents

MaxIT Consulting - Max Corporate Group

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

39 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

MaxIT Consulting - Max Corporate Group in San Francisco is seeking an AI Evaluation Engineer focused on reinforcement learning and agents to design environments, evaluation systems, and the infrastructure used to train and assess long-horizon enterprise AI workflows.

You will build high‑fidelity representations of complex software environments, implement graders and reward signals, and deliver production‑quality tools for rollouts, trajectory inspection, and robust evaluation.

Qualifications

  • Hands-on experience with AI environments, evaluations, reinforcement learning infrastructure, or related agent-training systems.
  • Strong software engineering fundamentals.
  • Demonstrated ability to build and ship technical infrastructure.
  • Understanding of evaluation methodology, reward design, graders, and agent trajectories.
  • Ability to work across languages and technology stacks based on system requirements.

Responsibilities

  • Design evaluation environments for long-horizon enterprise agent workflows.
  • Define tasks, state, tools, graders, and reward signals used to evaluate and improve agents.
  • Build high-fidelity representations of complex enterprise software environments.
  • Develop infrastructure for rollouts, orchestration, trajectory inspection, and grader pipelines.
  • Measure both correctness and efficiency across multi-step agent behavior.
  • Investigate evaluation failures, reward-quality issues, and agent behavior.
  • Build production-quality systems rather than notebook-only research prototypes.

Skills

AI environments
RL infrastructure
Software engineering
Evaluation methodology
Multi-language proficiency

Job description

San Francisco, California | Primarily On-site

We are seeking an AI Evaluation Engineer – Reinforcement Learning & Agents to build the environments, evaluation systems, and supporting infrastructure used to train and assess long-horizon enterprise AI agents.

The Opportunity

You will work on the engineering and research problems behind realistic agent environments, post-training systems, and reliable evaluation of complex multi-step workflows.

Key Responsibilities
  • Design evaluation environments for long-horizon enterprise agent workflows.
  • Define tasks, state, tools, graders, and reward signals used to evaluate and improve agents.
  • Build high-fidelity representations of complex enterprise software environments.
  • Develop infrastructure for rollouts, orchestration, trajectory inspection, and grader pipelines.
  • Measure both correctness and efficiency across multi-step agent behavior.
  • Investigate evaluation failures, reward-quality issues, and agent behavior.
  • Build production-quality systems rather than notebook-only research prototypes.
Required Qualifications
  • Hands‑on experience with AI environments, evaluations, reinforcement learning infrastructure, or related agent‑training systems.
  • Strong software engineering fundamentals.
  • Demonstrated ability to build and ship technical infrastructure.
  • Understanding of evaluation methodology, reward design, graders, and agent trajectories.
  • Ability to work across languages and technology stacks based on system requirements.
Candidate Profile

A PhD is not required. Strong engineering and shipped environment or evaluation systems are more important than academic credentials or publication history.

Seniority

The opportunity is open to exceptional new graduates, early‑career engineers, and experienced senior candidates. Selection is based primarily on engineering strength and relevant technical work.

Work Arrangement

The role is anchored in San Francisco with a strong preference for in‑person collaboration. Limited flexibility may be considered case by case for exceptional candidates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Agent Evaluation Infrastructure Engineer
Agent Evaluation Infrastructure Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 190,000
RL Environments Engineer
RL Environments Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Applied AI Deployment Engineer
Applied AI Deployment Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 110,000 - 160,000
AI Evaluation Engineer – RL Environments & Agents
AI Evaluation Engineer – RL Environments & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 150,000 - 210,000
Enterprise Agent Systems Engineer
Enterprise Agent Systems Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 150,000 - 230,000
Forward Deployed AI Engineer – Agent Systems
Forward Deployed AI Engineer – Agent Systems

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 120,000 - 170,000
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
AI Engineer, Evals & Agent Quality
AI Engineer, Evals & Agent Quality

Town.com, Inc. • San Francisco (CA)

On-site
USD 250,000 - 300,000
Remote Senior Software Engineer: AI Agent Evaluation
Remote Senior Software Engineer: AI Agent Evaluation

YO AI Labs • Washington

Remote
USD 60,000 - 90,000
Agent Engineer San Francisco CA
Agent Engineer San Francisco CA

AHU Technologies Inc • Washington

On-site
USD 120,000 - 150,000