Senior ML / Evaluation Engineer

EPAM Systems, Inc.

Portugal

Teletrabalho

EUR 70 000 - 110 000

Tempo integral

Há 10 dias
Gerador de candidaturas

Uma candidatura feita para esta oferta — um currículo e uma carta de apresentação personalizados que vão ao encontro do anúncio.

Ultrapassa os filtros ATS

Resumo da oferta

EPAM Systems, Inc. is seeking a Senior ML / Evaluation Engineer to design and implement advanced evaluation frameworks for an Enterprise Agent Development Platform. You will build LLM-as-judge evaluators and deterministic validators to enforce quality before release.

You will collaborate with platform, orchestration, and DevOps teams to ensure scalable, observable evaluation in production, leveraging AWS services, OpenTelemetry, and CloudWatch for monitoring and scoring.

Qualificações

  • 5+ years of ML engineering or AI evaluation experience.
  • Hands-on experience designing LLM evaluation frameworks.
  • Experience implementing CI/CD gates for ML QA.
  • Proficiency in Python for evaluation logic.

Responsabilidades

  • Design and implement multi-layer evaluation frameworks for agent workflows.
  • Build LLM-as-judge evaluators using AWS AgentCore modules.
  • Develop deterministic Lambda evaluators for rule-based checks.
  • Define enterprise evaluation standards with scoring criteria.
  • Integrate CI/CD deployment gates for automated pipelines.
  • Enable online evaluation in production with sampling strategies.

Conhecimentos

ML engineering
AI evaluation frameworks
LLM evaluation
CI/CD deployment
Python
OpenTelemetry
Observability
Cloud monitoring

Ferramentas

AWS AgentCore
AWS Bedrock
AWS Lambda
CloudWatch
OpenTelemetry

Descrição da oferta de emprego

We're looking for a Senior ML / Evaluation Engineer to join our team in Portugal in a fully remote working mode. In this role, you will own the design and implementation of advanced evaluation frameworks for an Enterprise Agent Development Platform—a production-grade, cloud-native ecosystem enabling scalable, secure AI agent deployment. You will create evaluation strategies that combine LLM-as-judge grading with deterministic checks, define enterprise evaluation standards, and implement CI/CD deployment gates to enforce quality metrics prior to release. This position requires strong expertise in ML system testing, evaluation design, and integration into automated pipelines for agentic environments.ResponsibilitiesDesign and implement multi-layer evaluation frameworks for agentic workflows and AI-driven applicationsBuild LLM-as-judge evaluators leveraging AWS AgentCore built-in modules and custom logic for correctness and helpfulness checksDevelop deterministic evaluators as AWS Lambda functions for rule-based validationDefine enterprise evaluation standards, including mandatory dimensions, scoring criteria, and pass/fail thresholdsImplement CI/CD deployment gates using on-demand evaluation modes to enforce quality in automated pipelinesEnable online evaluation in production by integrating sampling-based evaluation strategies and PII detection guardrailsIncorporate observability signals (OpenTelemetry spans) from AWS AgentCore into grading frameworks for trace-level assessmentGenerate metrics, logs, and dashboards from evaluation outcomes via CloudWatch or equivalent monitoring platformsCollaborate with platform, orchestration, and DevOps teams to maintain evaluation reliability and scalabilityRequirements5+ years of experience in ML engineering, AI evaluation frameworks, or AI platform developmentHands-on expertise designing LLM evaluation frameworks (LLM-as-judge and deterministic graders)Practical experience implementing CI/CD deployment gates for ML model or AI agent quality assuranceProficiency in Python for building evaluation logic (deterministic Lambda-based evaluators)Strong understanding of advanced validation dimensions, including multi-turn context integrity and workflow-level scoringNice to haveFamiliarity with AWS AgentCore Evaluations API (CreateEvaluation, GetEvaluationResult)Exposure to AWS Bedrock Guardrails for compliance and sensitive data validationExperience integrating evaluation metrics into AWS CloudWatch for monitoring and alertingKnowledge of OTel instrumentation and trace ingestion for quality scoring inputsWe offerCompetitive compensation depending on experience and skillsVariety of projects within one companyBeing a part of a project following engineering excellence standardsIndividual career path and professional growth opportunitiesInternal events and communitiesFlexible work hours
Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

Agent Evaluation Engineer — Build-Time Framework & Deployment Gates
Agent Evaluation Engineer — Build-Time Framework & Deployment Gates

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 70 000 - 100 000
Senior ML Evaluation Engineer (Remote) — CI/CD QA
Senior ML Evaluation Engineer (Remote) — CI/CD QA

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 70 000 - 110 000
Remote Agent Evaluation Engineer: Build-Time Frameworks
Remote Agent Evaluation Engineer: Build-Time Frameworks

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 70 000 - 100 000
Senior AI Engineer (LLM/Agentic Systems) - Full Remote Portugal
Senior AI Engineer (LLM/Agentic Systems) - Full Remote Portugal

HumanIT Digital Consulting • Porto

Teletrabalho
EUR 24 552 - 31 248
15th month salary
Health insurance for family
Birthday off
+2
Senior AI Engineer
Senior AI Engineer

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 90 000 - 130 000
Flexible work hours
Internal events and communities
Observability Engineer — AgentCore Telemetry
Observability Engineer — AgentCore Telemetry

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 70 000 - 110 000
Flexible work hours
Competitive compensation
Data AI Solution Architect
Data AI Solution Architect

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 90 000 - 140 000
Competitive compensation
Flexible work hours
Internal events and communities
Senior Platform Engineer — Agent Gateway Lead
Senior Platform Engineer — Agent Gateway Lead

EPAM Systems, Inc. • Portugal

Teletrabalho
EUR 90 000 - 135 000
Competitive compensation
Flexible work hours
Internal events and communities
+1
Senior / Lead Machine Learning Engineer (AI & LLM Systems) - Hybrid (2 days office)
Senior / Lead Machine Learning Engineer (AI & LLM Systems) - Hybrid (2 days office)

HumanIT Digital Consulting • Lisboa

Presencial
EUR 70 000 - 100 000
Hybrid work model
Opportunity for technical leadership
Mentorship opportunities
Lead Engineer, AI Platform
Lead Engineer, AI Platform

Lever, Inc. • Portugal

Teletrabalho
EUR 136 000 - 166 000
Fully remote environment
Equity & growth potential
Flexible schedule & autonomy
+1