Senior QA / ML Tester

EPAM Systems

Poland

Hybrid

PLN 180,000 - 260,000

Full time

14 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Hybrid work design
Work remotely within Poland
Opportunity to work abroad up to 60/90

Job summary

EPAM Systems is seeking a Senior QA / ML Tester to own quality assurance for the Agent Evaluation Framework built on AWS AgentCore. You will design, implement, and maintain a functional test suite for AI/ML evaluation pipelines, ensuring reliability in an Agile team.

The role requires hands-on testing across on-demand and online modes, OpenTelemetry integration validation, and defect reporting with Jira. Hybrid work in Poland is available, with potential for travel and collaboration with

Qualifications

  • 5+ years of production QA automation or ML/AI system testing.
  • Experience testing AI/LLM evaluation pipelines and agent behavior.
  • Proficiency in Python test automation with pytest, fixtures, parametrize, and mocking.
  • Knowledge of AWS AgentCore Evaluation API for on-demand and online modes.
  • Familiarity with OpenTelemetry/ADOT for trace-based testing.
  • REST API testing, including request/response validation and authentication.
  • CI/CD integration using GitHub Actions/Jenkins.
  • Test data management with known-good/known-bad session pairs.
  • Defect lifecycle management with Jira.
  • Understanding of LLM/agent evaluation concepts and evaluation contracts.
  • English at least B2+ for collaboration.

Responsibilities

  • Design and implement a functional test suite for AWS AgentCore Evaluation API using pytest.
  • Develop and maintain integration tests for on-demand evaluation, CI/CD integrated.
  • Validate online mode sampling accuracy; design scenarios and acceptance criteria.
  • Assess non-AgentCore runtime evaluation feasibility and deliver findings.
  • Test OpenTelemetry trace-based evaluation inputs and ADOT trace ingestion.
  • Collaborate with platform engineers to clarify evaluation contracts and reproduce defects.
  • Maintain test documentation in Confluence/Jira.
  • Participate in Agile ceremonies and EngX practices such as code reviews.
  • Contribute to test coverage reporting and CI/CD improvements.

Skills

Python testing
pytest
REST API testing
CI/CD integration
OpenTelemetry testing
AWS AgentCore
ML/AI system testing
Jira bug reports

Tools

GitHub Actions
Jenkins
unittest.mock
moto
OpenTelemetry

Job description

We are looking for a Senior QA / ML Tester to join the AI Platform team and take ownership of quality assurance for the Agent Evaluation Framework built on AWS AgentCore. This role involves designing, implementing, and maintaining a functional test suite that validates the correctness of AI agent evaluation pipelines — covering on-demand integration testing, online sampling accuracy, and multi-evaluator execution — and delivering a feasibility assessment for non-AgentCore runtime evaluation scenarios. This is a hands‑on, production‑focused role at the intersection of software quality engineering and AI/ML system testing, operating within an Agile delivery team and contributing to the reliability of enterprise‑grade agentic AI infrastructure.

Responsibilities
  • Design and implement a functional test suite for the AWS AgentCore Evaluation API using pytest, covering known-good / known-bad session pair validation, multi-evaluator execution correctness, and edge case handling
  • Develop and maintain integration tests for on-demand evaluation mode, integrated into the CI/CD pipeline with automated execution on each build
  • Validate online mode sampling accuracy, design test scenarios, define acceptance criteria, and report deviations with reproducible evidence
  • Conduct and document a feasibility assessment for non-AgentCore runtime evaluation: analyze alternative runtimes, define evaluation methodology, and deliver a structured findings report
  • Test OpenTelemetry trace-based evaluation inputs and validate ADOT trace ingestion, trace structure correctness, and evaluator input integrity
  • Collaborate with platform engineers to clarify evaluation contracts, reproduce defects, and align on quality gates
  • Maintain test documentation, including test plans, test reports, defect logs, and evaluation feasibility artifacts in Confluence/Jira
  • Participate in Agile ceremonies, including sprint planning, daily standups, demos, and retrospectives
  • Contribute to EngX practices such as code review of test scripts, CI/CD pipeline integration, and test coverage reporting
Requirements
  • 5+ years of production experience in QA automation or ML/AI system testing
  • Proven experience testing AI/LLM systems, including evaluation pipelines, model outputs, or agent behavior validation
  • Proficiency in Python test automation, including pytest, fixtures, parametrize, and mocking (unittest.mock, moto)
  • Knowledge of AWS AgentCore Evaluation API, covering on-demand and online evaluation modes
  • Familiarity with OpenTelemetry / ADOT for trace-based evaluation input testing and trace structure validation
  • Skills in REST API testing, including request/response validation and authentication (SigV4, bearer tokens)
  • Experience with CI/CD integration using GitHub Actions, Jenkins, or equivalent, including test pipeline configuration
  • Background in test data management, including known-good / known-bad session pair design and synthetic trace generation
  • Expertise in functional and integration test design for AI/ML evaluation pipelines
  • Competency in defect lifecycle management, including Jira, reproducible bug reports, and root cause analysis
  • Understanding of LLM/agent evaluation concepts, such as correctness scoring, sampling strategies, and evaluator chaining
  • Ability to work independently after onboarding, manage own tasks, report status, and elevate blockers proactively
  • Strong analytical skills to define test scenarios from ambiguous or evolving specifications
  • English B2+ level, written and verbal, for daily collaboration with distributed teams
Nice to have
  • Hands‑on experience with AWS AgentCore Evaluation API or AWS Bedrock testing
  • Experience testing OpenTelemetry / distributed tracing pipelines
  • Familiarity with multi-evaluator execution patterns and correctness validation strategies
  • Experience writing feasibility assessments or technical reports for stakeholders
  • Knowledge of agentic AI frameworks (LangGraph, Strands Agents) sufficient to understand evaluation contracts
  • ISTQB CT-AI certification or equivalent AI testing qualification
  • Experience with AI Ready / AI Practitioner practices at EPAM (prompt engineering, AI-assisted test design)
We offer
  • We gather like-minded people:
    • Top tech minds driving innovation in AI, cloud and digital platform modernization
    • Supportive team and agile, startup-like culture
    • Hybrid by design mode and opportunity to work remotely within Poland
    • Chance to work abroad for up to 60 days annually
    • Business-driven relocation opportunities
  • We provide growth opportunities:
    • Career development programs
    • Thought leadership, mentoring, soft skills and well-being programs
    • Certification (Anthropic, Gemini, GCP, Azure, AWS)
    • English classes
  • We cover it all:
    • Stable pay
    • Participation in the Employee Stock Purchase Plan with a 15% discount
    • Benefits package (health insurance, multisport, shopping vouchers)
    • Referral bonuses up to $2,000
    • Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and more
    • Corporate, social and well-being events
  • Please, note:
    • Benefits listed above are available to employees only
    • We are open for working with Contractors. Terms of B2B cooperation agreements are agreed individually
    • We will reach out to selected candidates exclusively

EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior QA Engineer – ML/AI Evaluation Pipelines
Senior QA Engineer – ML/AI Evaluation Pipelines

EPAM Systems • Poland

Hybrid
PLN 180,000 - 260,000
Hybrid work design
Work remotely within Poland
Opportunity to work abroad up to 60/90
Senior Security & Test Engineer - A2A
Senior Security & Test Engineer - A2A

EPAM Systems • Łódź

Hybrid
PLN 200,000 - 280,000
Health insurance
Multisport card
Relocation opportunities
+1
AI Architect
AI Architect

EPAM Systems • Kraków

On-site
PLN 250,000 - 380,000
Hybrid-friendly work patterns
Relocation opportunities
English classes
+1
Senior AI & Agentic Systems Engineer (Java)
Senior AI & Agentic Systems Engineer (Java)

EPAM Systems • Wrocław

Hybrid
PLN 210,000 - 320,000
Hybrid work model
Work abroad up to 60 days annually
Relocation opportunities
+2
AI Architect
AI Architect

EPAM Systems • Województwo pomorskie

Hybrid
PLN 240,000 - 360,000
Hybrid by design—remote options in POL
Relocation opportunities within Poland
Certification programs (Anthropic, GCP
+1
Senior AI/ML Solution Engineer (AI Agentic Security Testing)
Senior AI/ML Solution Engineer (AI Agentic Security Testing)

EPAM Systems • Łódź

Hybrid
PLN 250,000 - 390,000
Hybrid work model
Relocation opportunities
English classes
+2
Senior AI-Native Engineer
Senior AI-Native Engineer

EPAM Systems • Kraków

Hybrid
PLN 200,000 - 380,000
Hybrid by design
Relocation opportunities
Employee Stock Purchase Plan
+3
Senior Software Quality Engineer (Test Automation)
Senior Software Quality Engineer (Test Automation)

Tenarai Europe • Kraków

Hybrid
PLN 240,000 - 360,000
Hybrid work model
Onsite parking space for employees
Life Insurance
+6
Lead .NET Engineer (Full-Stack, Front-End, Cloud & AI)
Lead .NET Engineer (Full-Stack, Front-End, Cloud & AI)

EPAM Systems • Poznań

Hybrid
PLN 280,000 - 420,000
Hybrid by design
Work remotely within Poland
Relocation opportunities
+1
Lead .NET Engineer (Full-Stack, Front-End, Cloud & AI)
Lead .NET Engineer (Full-Stack, Front-End, Cloud & AI)

EPAM Systems • Województwo pomorskie

Hybrid
PLN 240,000 - 360,000
Hybrid by design
Remote within Poland
Career development programs
+2