Agent Evaluation Infrastructure Engineer

MaxIT Consulting - Max Corporate Group

San Francisco (CA)

On-site

USD 160,000 - 230,000

Full time

23 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

MaxIT Consulting - Max Corporate Group in San Francisco, CA seeks an Agent Evaluation Infrastructure Engineer to build environments, evaluation systems and supporting infrastructure for training and assessing long-horizon enterprise AI agents. The role anchors on-site with emphasis on real-world production quality rather than notebook-style research.

You will design evaluation environments, define tasks, state, tools, graders and reward signals, and develop scalable pipelines for rollouts and

Qualifications

  • Hands-on experience with AI environments, evaluations, reinforcement learning infrastructure, or related agent-training systems.
  • Strong software engineering fundamentals.
  • Demonstrated ability to build and ship technical infrastructure.
  • Understanding of evaluation methodology, reward design, graders, and agent trajectories.
  • Ability to work across languages and technology stacks based on system requirements.

Responsibilities

  • Design evaluation environments for long-horizon enterprise agent workflows.
  • Define tasks, state, tools, graders, and reward signals used to evaluate and improve agents.
  • Build high-fidelity representations of complex enterprise software environments.
  • Develop infrastructure for rollouts, orchestration, trajectory inspection, and grader pipelines.
  • Measure both correctness and efficiency across multi-step agent behavior.
  • Investigate evaluation failures, reward-quality issues, and agent behavior.
  • Build production-quality systems rather than notebook-only research prototypes.

Skills

AI environments
RL infra
Software engineering
Evaluation methods
Cross-language

Tools

Grader pipelines
Evaluation tooling

Job description

San Francisco, California | Primarily On-site


We are seeking an Agent Evaluation Infrastructure Engineer to build the environments, evaluation systems, and supporting infrastructure used to train and assess long-horizon enterprise AI agents.


The Opportunity

You will work on the engineering and research problems behind realistic agent environments, post-training systems, and reliable evaluation of complex multi-step workflows.


Key Responsibilities


  • Design evaluation environments for long-horizon enterprise agent workflows.

  • Define tasks, state, tools, graders, and reward signals used to evaluate and improve agents.

  • Build high-fidelity representations of complex enterprise software environments.

  • Develop infrastructure for rollouts, orchestration, trajectory inspection, and grader pipelines.

  • Measure both correctness and efficiency across multi-step agent behavior.

  • Investigate evaluation failures, reward-quality issues, and agent behavior.

  • Build production-quality systems rather than notebook-only research prototypes.


Required Qualifications


  • Hands-on experience with AI environments, evaluations, reinforcement learning infrastructure, or related agent-training systems.

  • Strong software engineering fundamentals.

  • Demonstrated ability to build and ship technical infrastructure.

  • Understanding of evaluation methodology, reward design, graders, and agent trajectories.

  • Ability to work across languages and technology stacks based on system requirements.


Candidate Profile

A PhD is not required. Strong engineering and shipped environment or evaluation systems are more important than academic credentials or publication history.


Seniority

The opportunity is open to exceptional new graduates, early-career engineers, and experienced senior candidates. Selection is based primarily on engineering strength and relevant technical work.


Work Arrangement

The role is anchored in San Francisco with a strong preference for in-person collaboration. Limited flexibility may be considered case by case for exceptional candidates.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Enterprise AI Evaluation Environments Engineer
Enterprise AI Evaluation Environments Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
RL Environments Engineer
RL Environments Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI Evaluation Engineer – Reinforcement Learning & Agents
AI Evaluation Engineer – Reinforcement Learning & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Research Engineer – RL Infrastructure & Agent Environments
Research Engineer – RL Infrastructure & Agent Environments

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 150,000 - 190,000
Enterprise AI Evaluation Infra Engineer — SF On-Site
Enterprise AI Evaluation Infra Engineer — SF On-Site

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 160,000 - 230,000
Forward Deployed AI Engineer – Agent Systems
Forward Deployed AI Engineer – Agent Systems

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 170,000 - 250,000
Enterprise Agent Systems Engineer
Enterprise Agent Systems Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 190,000
AI Evaluation Engineer: RL Environments & Agents
AI Evaluation Engineer: RL Environments & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Applied AI Deployment Engineer
Applied AI Deployment Engineer

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 120,000 - 170,000
Forward Deployed AI Engineer – Agents & Research Systems
Forward Deployed AI Engineer – Agents & Research Systems

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000