Join the Enterprise Agent Development Platform project at EPAM. We are building a cloud-native platform that enables engineering teams to develop, deploy, and operate AI agents in production.The platform combines modern agent frameworks, AWS infrastructure, CI/CD, observability, and governance to make AI development faster and more reliable.As a Python AI Evaluation Engineer, you will own a key part of this platform: building the capabilities that measure and validate the quality of AI solutions. You will design evaluation approaches, develop custom evaluators, and integrate quality checks into the software delivery lifecycle.This role is a strong fit for an engineer who enjoys solving new GenAI challenges and turning them into practical, automated engineering solutions.ResponsibilitiesDesign and implement evaluation frameworks for LLM-based applications and AI agentsDevelop LLM-as-a-Judge and deterministic, code-based evaluatorsBuild custom Python evaluators for quality and behavioral checksDefine evaluation criteria, metrics, thresholds, and acceptance rulesEvaluate agent behavior across individual responses, tool calls, and complete workflowsWork with OpenTelemetry traces and spans as evaluation dataIntegrate evaluations into CI/CD pipelines and automated deployment gatesEnable continuous quality monitoring of solutions in productionEstablish reusable evaluation patterns and engineering standardsWork closely with AI Engineers, Architects, and Platform Engineers to embed quality into the development processRequirements5+ years of experience in ML Engineering, AI Engineering, or AI Platform EngineeringStrong Python development experienceHands-on experience with LLM/GenAI evaluationExperience designing and implementing evaluation frameworksExperience developing custom or deterministic evaluatorsExperience integrating AI/ML quality checks into CI/CDGood understanding of LLM and AI agent architecturesNice to haveHands-on experience with AWS AgentCore EvaluationExperience with AWS Bedrock Guardrails, including PII detectionKnowledge of CloudWatch metrics and production monitoringExperience with OpenTelemetryFamiliarity with LangGraph, Strands Agents, or similar agent frameworksWe offerCONTINUOUS UPSKILLING, LEARNING & DEVELOPMENTDiversity of tasks and projectsAssessment center for objective review of competency levelPersonal development planMentoring programs and leadership developmentCertification and professional development supportAccess to learning platforms including more than 2,500 internal coursesEnglish courses taught by certified teachersCORPORATE BENEFITSExtra leave daysReferral bonusesCOMPENSATION PACKAGECompetitive compensation paid in USDRegular salary and performance reviewsMEDICAL & HEALTHCAREPrivate health insuranceWell-being eventsWORKING ENVIRONMENTRecreation areas and kitchensTea, coffee and snacksSports equipment and game consolesIT EquipmentMicrosoft’s Software Assurance Home Use Program (HUP)Please note that our Talent Attraction Team reviews applications and CVs submitted in English.EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.