Agent Evaluation Engineer — Build-Time Framework & Deployment Gates

EPAM Systems Inc

United States

Remote

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

EPAM Systems is seeking an Agent Evaluation Engineer to design and maintain an evaluation framework for AI agents, including automated tests and CI/CD deployment gates. The role focuses on creating multi-layer evaluation suites that blend deterministic checks with LLM-powered graders, simulating multi-turn conversations, and defining reliability metrics.

You will implement staging validations, shadow-mode traffic analysis, and A/B rollout strategies, with feedback loops from production

Qualifications

  • 4+ years building automated testing or evaluation frameworks for ML/LLM/agentic systems.
  • Experience designing multi-layer evaluation suites with deterministic and LLM-based graders.
  • Expertise with CI/CD pipelines and metric-based quality gates for automated deployments.
  • Familiarity with LangGraph or similar agent orchestration frameworks.
  • Experience with shadow-mode, canary, or A/B deployment practices for ML platforms.

Responsibilities

  • Design and implement build-time evaluation frameworks for agentic workflows using LangGraph or similar orchestration.
  • Create deterministic and LLM-as-judge grading pipelines covering reasoning, trajectory accuracy, and output quality.
  • Develop test harnesses for multi-turn conversational simulations and context-retention scoring.
  • Define reliability assessment methods including multi-trial metrics (pass@k, pass^k).
  • Implement CI/CD deployment gates that enforce quality thresholds and block releases not meeting standards.
  • Integrate staging validation, shadow-mode traffic comparison, and A/B rollout control in deployment pipelines.
  • Leverage AWS AgentCore Evaluations API for on-demand and online scoring components connected to production feedback.
  • Convert production incidents into reusable regression cases for continuous quality improvement.
  • Collaborate with engineering and DevOps teams to embed evaluation gates into automated workflows.

Skills

Automated testing frameworks
Evaluation frameworks
CI/CD pipelines
LangGraph orchestration
LLM-based graders
AWS AgentCore Evaluations
Shadow-mode/canary/A/B deployment
Multi-turn conversation simulations

Tools

LangGraph
AWS AgentCore Evaluations API

Job description

We're looking for an Agent Evaluation Engineer — Build-Time Framework & Deployment Gates to join our team in Portugal in a fully remote working mode. In this role, you will design and maintain an evaluation framework for AI agents, ensuring quality and compliance through automated tests and CI/CD deployment gates. You will develop multi-layer evaluation suites that blend deterministic checks with LLM-powered graders, simulate multi-turn conversations, and define reliability metrics. The position also involves implementing staging validations, shadow-mode traffic analysis, and A/B rollout strategies, with feedback loops from production environments to enhance overall system robustness.

Responsibilities
  • Design and implement build-time evaluation frameworks for agentic workflows using LangGraph or comparable orchestration frameworks
  • Create deterministic and LLM-as-judge grading pipelines covering reasoning, trajectory accuracy, and output quality
  • Develop test harnesses for multi-turn conversational simulations and context-retention scoring
  • Define reliability assessment methods including multi-trial metrics (pass@k, pass^k)
  • Implement CI/CD deployment gates that enforce quality thresholds and block releases not meeting standards
  • Integrate staging validation, shadow-mode traffic comparison, and A/B rollout control in deployment pipelines
  • Leverage AWS AgentCore Evaluations for on-demand and online scoring components connected to production feedback
  • Convert production incidents into reusable regression cases for continuous quality improvement
  • Collaborate with engineering and DevOps teams to embed evaluation gates into automated workflows
Requirements
  • 4+ years of experience in building automated testing or evaluation frameworks for ML, LLM, or agentic systems
  • Proven hands‑on experience designing multi-layer evaluation suites with deterministic and LLM-based graders
  • Expertise with CI/CD pipelines and implementing metric-based quality gates for automated deployments
  • Practical knowledge of LangGraph or similar agent orchestration frameworks
  • Strong background in designing simulation-based evaluation strategies and conversation-level tests
  • Nice to have Experience with AWS AgentCore Evaluations API (CreateEvaluation, custom evaluators)
  • Familiarity with shadow-mode, canary, or A/B deployment practices for ML-based platforms
  • Background in transforming production failures into build-time regression tests for agent workflows
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Agent Evaluation Engineer
Remote AI Agent Evaluation Engineer

EPAM Systems Inc • United States

Remote
USD 140,000 - 190,000
Staff Software Engineer, Agent Eval Platform
Staff Software Engineer, Agent Eval Platform

Servicenow • Santa Clara (CA)

On-site
USD 180,000 - 320,000
Data Scientist, Agent Evaluations & Quality
Data Scientist, Agent Evaluations & Quality

Clera • Palo Alto (CA)

On-site
USD 150,000 - 210,000
QA Engineer - Agentic Systems
QA Engineer - Agentic Systems

Meet Life Sciences • New York (NY)

On-site
USD 110,000 - 170,000
QA AI Automation Engineer
QA AI Automation Engineer

Dynasty Financial Partners • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

On-site
USD 162,000 - 198,000
QA AI Automation Engineer
QA AI Automation Engineer

Dynasty Financial Partners, LLC • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
QA / Automation Engineer Agentic AI
QA / Automation Engineer Agentic AI

Compunnel, Inc. • Atlanta (GA), Northern (KY)

Hybrid
USD 110,000 - 160,000
Software Engineering Director, Agentic Evaluations (Fully Remote)
Software Engineering Director, Agentic Evaluations (Fully Remote)

Partner Company • United States

Remote
USD 260,000 - 320,000
Unlimited PTO
Equity
Parental leave
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Austin (TX)

On-site
USD 68,191 - 97,120
Competitive pay
Flexible schedule
Experience in advanced AI projects