Agent Evaluation Engineer — Build-Time Framework & Deployment Gates

EPAM Systems, Inc.

Portugal

Teletrabalho

EUR 70 000 - 100 000

Tempo integral

Há 9 dias
Gerador de candidaturas

Transforma esta função numa entrevista — um currículo e uma carta de apresentação criados à volta do que este empregador procura.

Ultrapassa os filtros ATS

Resumo da oferta

EPAM Systems, Inc. is seeking an Agent Evaluation Engineer to build evaluation frameworks for AI agents, combining deterministic and LLM-based grading with CI/CD deployment gates.

You will design multi-layer evaluation suites, simulate multi-turn conversations, and establish reliability metrics, including pass@k metrics and staging validations. The role is fully remote in Portugal, offering exposure to diverse projects and strong collaboration with DevOps to embed gates into automated pipelines,

Qualificações

  • 4+ years of experience building automated testing or evaluation frameworks for ML, LLM, or agentic systems.
  • Proven hands-on experience designing multi-layer evaluation suites with deterministic and LLM-based graders.
  • Expertise with CI/CD pipelines and implementing metric-based quality gates for automated deployments.
  • Practical knowledge of LangGraph or similar agent orchestration frameworks.

Responsabilidades

  • Design and implement build-time evaluation frameworks for agentic workflows.
  • Create deterministic and LLM-as-judge grading pipelines covering reasoning, trajectory accuracy, and output quality.
  • Develop test harnesses for multi-turn conversational simulations and context-retention scoring.
  • Define reliability assessment methods including multi-trial metrics (pass@k, pass^k).
  • Implement CI/CD deployment gates that enforce quality thresholds and block releases not meeting standards.
  • Integrate staging validation, shadow-mode traffic comparison, and A/B rollout control in deployment pipelines.
  • Leverage AWS AgentCore Evaluations for on-demand and online scoring components connected to production feedback.
  • Convert production incidents into reusable regression cases for continuous quality improvement.

Conhecimentos

CI/CD pipelines
Evaluation frameworks
LangGraph
Deterministic + LLM grading
Multi-turn tests
Simulation testing
Reliability metrics
A/B deployment

Ferramentas

LangGraph
AWS AgentCore
AWS AgentCore API
AgentCore Evaluations

Descrição da oferta de emprego

We're looking for an Agent Evaluation Engineer — Build-Time Framework & Deployment Gates to join our team in Portugal in a fully remote working mode. In this role, you will design and maintain an evaluation framework for AI agents, ensuring quality and compliance through automated tests and CI/CD deployment gates. You will develop multi-layer evaluation suites that blend deterministic checks with LLM-powered graders, simulate multi-turn conversations, and define reliability metrics. The position also involves implementing staging validations, shadow-mode traffic analysis, and A/B rollout strategies, with feedback loops from production environments to enhance overall system robustness.ResponsibilitiesDesign and implement build-time evaluation frameworks for agentic workflows using LangGraph or comparable orchestration frameworksCreate deterministic and LLM-as-judge grading pipelines covering reasoning, trajectory accuracy, and output qualityDevelop test harnesses for multi-turn conversational simulations and context-retention scoringDefine reliability assessment methods including multi-trial metrics (pass@k, pass^k)Implement CI/CD deployment gates that enforce quality thresholds and block releases not meeting standardsIntegrate staging validation, shadow-mode traffic comparison, and A/B rollout control in deployment pipelinesLeverage AWS AgentCore Evaluations for on-demand and online scoring components connected to production feedbackConvert production incidents into reusable regression cases for continuous quality improvementCollaborate with engineering and DevOps teams to embed evaluation gates into automated workflowsRequirements4+ years of experience in building automated testing or evaluation frameworks for ML, LLM, or agentic systemsProven hands-on experience designing multi-layer evaluation suites with deterministic and LLM-based gradersExpertise with CI/CD pipelines and implementing metric-based quality gates for automated deploymentsPractical knowledge of LangGraph or similar agent orchestration frameworksStrong background in designing simulation-based evaluation strategies and conversation-level testsNice to haveExperience with AWS AgentCore Evaluations API (CreateEvaluation, custom evaluators)Familiarity with shadow-mode, canary, or A/B deployment practices for ML-based platformsBackground in transforming production failures into build-time regression tests for agent workflowsWe offerCompetitive compensation depending on experience and skillsVariety of projects within one companyBeing a part of a project following engineering excellence standardsIndividual career path and professional growth opportunitiesInternal events and communitiesFlexible work hours
Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 70 000 - 110 000
Remote Agent Evaluation Engineer: Build-Time Frameworks
Remote Agent Evaluation Engineer: Build-Time Frameworks

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 70 000 - 100 000
Data AI Solution Architect
Data AI Solution Architect

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 90 000 - 140 000
Competitive compensation
Flexible work hours
Internal events and communities
Senior ML Evaluation Engineer (Remote) — CI/CD QA
Senior ML Evaluation Engineer (Remote) — CI/CD QA

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 70 000 - 110 000
Senior Platform Engineer — Agent Gateway Lead
Senior Platform Engineer — Agent Gateway Lead

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 90 000 - 135 000
Competitive compensation
Flexible work hours
Internal events and communities
+1
Remote LLM Engineer: Build Production AI Agents
Remote LLM Engineer: Build Production AI Agents

Xelerate • Porto

Híbrido
EUR 60 000 - 90 000
Remote work
Strong benefits
Work-life balance
+2
Observability Engineer — AgentCore Telemetry
Observability Engineer — AgentCore Telemetry

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 70 000 - 110 000
Flexible work hours
Competitive compensation
Senior AI Engineer
Senior AI Engineer

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 90 000 - 130 000
Flexible work hours
Internal events and communities
Senior AI Engineer (LLM/Agentic Systems) - Full Remote Portugal
Senior AI Engineer (LLM/Agentic Systems) - Full Remote Portugal

HumanIT Digital Consulting • Porto

Teletrabalho
EUR 24 552 - 31 248
15th month salary
Health insurance for family
Birthday off
+2
Backend Engineer — A2A Integration
Backend Engineer — A2A Integration

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 65 000 - 86 000
Remote work
Flexible hours