Remote AI Agent Evaluation Engineer

EPAM Systems Inc

United States

Remote

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

EPAM Systems is seeking an Agent Evaluation Engineer to design and maintain an evaluation framework for AI agents, including automated tests and CI/CD deployment gates. The role focuses on creating multi-layer evaluation suites that blend deterministic checks with LLM-powered graders, simulating multi-turn conversations, and defining reliability metrics.

You will implement staging validations, shadow-mode traffic analysis, and A/B rollout strategies, with feedback loops from production

Qualifications

  • 4+ years building automated testing or evaluation frameworks for ML/LLM/agentic systems.
  • Experience designing multi-layer evaluation suites with deterministic and LLM-based graders.
  • Expertise with CI/CD pipelines and metric-based quality gates for automated deployments.
  • Familiarity with LangGraph or similar agent orchestration frameworks.
  • Experience with shadow-mode, canary, or A/B deployment practices for ML platforms.

Responsibilities

  • Design and implement build-time evaluation frameworks for agentic workflows using LangGraph or similar orchestration.
  • Create deterministic and LLM-as-judge grading pipelines covering reasoning, trajectory accuracy, and output quality.
  • Develop test harnesses for multi-turn conversational simulations and context-retention scoring.
  • Define reliability assessment methods including multi-trial metrics (pass@k, pass^k).
  • Implement CI/CD deployment gates that enforce quality thresholds and block releases not meeting standards.
  • Integrate staging validation, shadow-mode traffic comparison, and A/B rollout control in deployment pipelines.
  • Leverage AWS AgentCore Evaluations API for on-demand and online scoring components connected to production feedback.
  • Convert production incidents into reusable regression cases for continuous quality improvement.
  • Collaborate with engineering and DevOps teams to embed evaluation gates into automated workflows.

Skills

Automated testing frameworks
Evaluation frameworks
CI/CD pipelines
LangGraph orchestration
LLM-based graders
AWS AgentCore Evaluations
Shadow-mode/canary/A/B deployment
Multi-turn conversation simulations

Tools

LangGraph
AWS AgentCore Evaluations API

Job description

EPAM Systems is seeking an Agent Evaluation Engineer to design and maintain an evaluation framework for AI agents, including automated tests and CI/CD deployment gates. The role focuses on creating multi-layer evaluation suites that blend deterministic checks with LLM-powered graders, simulating multi-turn conversations, and defining reliability metrics.

You will implement staging validations, shadow-mode traffic analysis, and A/B rollout strategies, with feedback loops from production

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Agent Engineer: Production Systems & Orchestration
Senior AI Agent Engineer: Production Systems & Orchestration

EPAM Systems Inc • United States

Remote
USD 150,000 - 210,000
Agent Evaluation Engineer — Build-Time Framework & Deployment Gates
Agent Evaluation Engineer — Build-Time Framework & Deployment Gates

EPAM Systems Inc • United States

Remote
USD 140,000 - 190,000
Remote AI Agent Evaluation Engineer
Remote AI Agent Evaluation Engineer

YO AI Labs • Los Angeles (CA)

Remote
USD 83,000 - 138,000
AI Evaluation Engineer: RL Environments & Agents
AI Evaluation Engineer: RL Environments & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Remote Senior Software Engineer: AI Agent Evaluation
Remote Senior Software Engineer: AI Agent Evaluation

YO AI Labs • Washington

Remote
USD 60,000 - 90,000
Remote AI Evaluation & Quality Assurance Engineer
Remote AI Evaluation & Quality Assurance Engineer

Equiliem • United States

Remote
USD 69,000 - 74,000
Remote AI Agent Evaluation Engineer
Remote AI Agent Evaluation Engineer

YO AI Labs • New York (NY)

Remote
USD 46,000 - 110,000
ML Engineer for AI Agents & Evaluation Frameworks
ML Engineer for AI Agents & Evaluation Frameworks

Enfint • United States

Remote
USD 140,000 - 190,000
Remote Senior AI Agent Evaluation Engineer
Remote Senior AI Agent Evaluation Engineer

YO AI Labs • San Francisco (CA)

Remote
USD 83,000 - 124,000
Senior AI Engineer - Agent Quality & Evaluations (Remote)
Senior AI Engineer - Agent Quality & Evaluations (Remote)

vibehackers • Northern (KY)

Hybrid
USD 150,000 - 250,000
Health insurance
Parental leave
Unlimited flexible time off
+2