Senior AI QA Test Automation Engineer – CX Agent

XM

Poland

On-site

PLN 180,000 - 280,000

Full time

30 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Attractive remuneration package
Intellectually stimulating work env.
International training opportunities

Job summary

XM is seeking a Senior AI QA Test Automation Engineer to lead technical QA for AI-driven CX Agent Suite, focusing on evaluation pipelines and multi-agent interactions in AWS Bedrock context.

Collaborate with QA leadership, Data Science, and Engineering to architect scalable evaluation systems, prototype agent orchestration, and integrate testing into GitLab CI/CD.

The role emphasizes hands-on automation, data-driven quality analytics, mentoring, and influencing engineering standards across teams.

Qualifications

  • 6+ years of hands-on QA/Test Automation experience.
  • 1+ years applying AI/ML in software testing.
  • Strong Python skills in AWS-centric environments; Java/TypeScript a plus.
  • Hands-on experience with LLM evaluation techniques (LLM-as-a-judge, human-in-the-loop).
  • Experience with DeepEval or similar evaluation frameworks; familiarity with agent-orchestration SDKs (Strands Agents).
  • Knowledge of LLM tooling, vector databases, and MLOps pipelines.
  • Experience integrating AI tooling into enterprise CI/CD (GitLab) and cloud-native environments (Docker, Kubernetes).
  • Strong communicator, able to influence technical decisions across teams.

Responsibilities

  • Design and build AI evaluation pipelines that assess multi-agent CX Agent Suite outputs for accuracy, relevance, tone, and safety.
  • Develop scalable AI/QA frameworks and tooling for the CX Agent Suite and beyond.
  • Evaluate and prototype Strands Agents or similar orchestration frameworks for self-evolving QA agents.
  • Integrate AI-driven testing into the existing GitLab CI/CD pipeline and cloud-native environments.
  • Lead research into AI testing methodologies and mentor engineers on code reviews and architectural guidance.

Skills

QA/Test Automation
Python
AI/ML in testing
CI/CD
Docker & Kubernetes
LLM evaluation
Strong communication

Education

BSc/MSc in Computer Science or AI

Tools

GitLab
Docker
Kubernetes
Terraform
OpenTelemetry
LangFuse
LangSmith
Arize

Job description

We're seeking a Senior AI QA Test Automation Engineer to take a leading technical role in AI quality and test automation. You will work primarily with our AWS Bedrock AgentCore-based CX Agent Suite, a multi-agent customer-support system covering intent routing, FAQs, deposit-status queries, and human handoff. You will design and build the evaluation and QA tooling required to validate AI agents reliably, while also contributing to technical standards and best practices across QA.

This is a Senior Individual Contributor (IC) role with a strong technical and strategic focus. You'll work closely with QA leadership, Data Science, and Engineering to architect scalable AI evaluation systems, establish technical standards, and drive the evolution of our AI evaluation framework. Our current evaluation harness is based on DeepEval and is evolving towards a broader agentic evaluation architecture. You will have the opportunity to evaluate and prototype emerging agent-orchestration approaches, including Strands Agents, and contribute to architectural decisions around their adoption.

The main responsibilities of the position include:
  • Act as the primary technical enabler for QA, building scalable AI/ML frameworks, libraries, and tooling for the CX Agent Suite and beyond
  • Design and build AI evaluation pipelines that assess the CX Agent Suite's multi-agent responses (Intent Detector, FAQ Agent, Missing-Deposits Agent) for accuracy, relevance, tone, hallucination rate, safety/guardrail compliance, and task completion — extending or replacing the current DeepEval-based evaluation harness
  • Evaluate and prototype Strands Agents (or comparable agent-orchestration frameworks) for building self-evolving, autonomous QA agents, and drive the go/no-go decision on adoption
  • Collaborate with QA, Data Science, and Engineering to integrate AI-driven testing into the existing GitLab CI/CD pipeline (dev → test → staging → prod), alongside Terraform-provisioned, EKS/AgentCore-hosted services
  • Build resilient and adaptive automation that can detect and respond to changes in agent behavior, Bedrock Guardrails configuration, and routing logic
  • Develop data-driven quality analytics, including root-cause analysis, quality trends, and intelligent test prioritization, leveraging existing observability and evaluation data from OpenTelemetry, CloudWatch, X-Ray traces, and LangFuse offline evaluation runs
  • Lead research into emerging AI testing methodologies and mentor engineers through code reviews, workshops, and architectural guidance
Main requirements:
  • BSc/MSc in Computer Science, AI, or a related discipline
  • 6+ years of hands‑on experience in QA/Test Automation, with strong experience designing and maintaining automation frameworks
  • 1+ years of experience applying AI/ML in software testing or QA process improvement, with a proven track record of bringing AI/ML solutions into production workflows
  • Strong Python skills, with experience working in Python/AWS-centric technology environments; Java and/or TypeScript is a plus
  • Hands‑on experience with LLM evaluation techniques, including LLM-as-a-judge, human‑in‑the‑loop evaluation, RAG, and multi‑agent orchestration patterns
  • Practical experience with DeepEval or comparable evaluation frameworks, along with familiarity with agent‑orchestration SDKs such as Strands Agents
  • Knowledge of LLM tooling, vector databases, and MLOps pipelines
  • Experience integrating AI tooling into enterprise CI/CD environments, particularly GitLab, and working with containerized cloud‑native environments such as Docker and Kubernetes/EKS
  • Strong communicator, able to influence technical decisions and drive engineering standards across teams
The following will be considered an advantage:
  • Experience with autonomous QA agents or agentic orchestration frameworks for self‑evolving test suites
  • Experience with LLM observability tools such as LangFuse, LangSmith, or Arize, particularly for measuring probabilistic and adversarial robustness
  • Knowledge of AI ethics, fairness, and bias detection for guardrail and model validation
  • Experience with gRPC, WebSockets, and/or HTTP/2
  • Experience with AWS Bedrock, plus familiarity with GCP Vertex AI and/or Azure AI
Benefit from:
  • Attractive remuneration package
  • Intellectually stimulating work environment
  • Continuous personal development and international training opportunities
The Hiring Experience: What Awaits You
  • Let’s Connect – Intro Chat with Talent Acquisition
  • Deep Dive – First Interview with Your Future Team
  • Final Connection – Final Interview

All applications will be treated with strict confidentiality!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior QA / ML Tester
Senior QA / ML Tester

EPAM Systems • Poland

Hybrid
PLN 180,000 - 260,000
Hybrid work design
Work remotely within Poland
Opportunity to work abroad up to 60/90
AI Quality Engineer
AI Quality Engineer

Globaldev Group • Warszawa

On-site
PLN 180,000 - 240,000
HR support
20 days vacation
Multisport
+2
Senior Automation Test Engineer – AI Systems (f/m/x)
Senior Automation Test Engineer – AI Systems (f/m/x)

Sii Poland • Katowice

On-site
PLN 180,000 - 260,000
Senior Automation Test Engineer – AI Systems (f/m/x)
Senior Automation Test Engineer – AI Systems (f/m/x)

Sii Poland • Poznań

On-site
PLN 180,000 - 280,000
Senior Automation Test Engineer – AI Systems (f/m/x)
Senior Automation Test Engineer – AI Systems (f/m/x)

Sii Poland • Toruń

On-site
PLN 180,000 - 240,000
Great Place to Work
Solid financial situation
Contracts with the biggest brands
+5
Senior Automation Test Engineer – AI Systems (f/m/x)
Senior Automation Test Engineer – AI Systems (f/m/x)

Sii Poland • Województwo pomorskie

On-site
PLN 180,000 - 240,000
Profit sharing
Medical care
Centre of internal trainings
+2
Senior Automation Test Engineer – AI Systems (f/m/x)
Senior Automation Test Engineer – AI Systems (f/m/x)

Sii Poland • Piła

On-site
PLN 180,000 - 240,000
Great Place to Work
Profit sharing
Center of internal trainings
+1
Senior Automation Test Engineer – AI Systems (f/m/x)
Senior Automation Test Engineer – AI Systems (f/m/x)

Sii Poland • Białystok

On-site
PLN 180,000 - 270,000
Great Place to Work
Centre of internal trainings
Profit sharing
+2
AI Quality Engineer
AI Quality Engineer

Start.io • Poland

On-site
PLN 180,000 - 270,000
Senior Automation Test Engineer – AI Systems (f/m/x)
Senior Automation Test Engineer – AI Systems (f/m/x)

Sii Poland • Wrocław

On-site
PLN 180,000 - 240,000
Great Place to Work
Solid financial situation
Centre of internal trainings
+2