Quality Assurance Tester

Intellias

Portugal

Presencial

EUR 45 000 - 65 000

Tempo integral

14 dias+

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Resumo da oferta

Intellias is seeking a QA / ML Tester – Evaluation Framework to ensure quality and reliability of AI evaluation systems across enterprise platforms. You will design validation strategies, test evaluator behavior across varied scenarios, and verify automated quality assessment pipelines in collaboration with AI engineers and platform teams.

The role sits in Intellias Core Architecture Team, contributing to a scalable, secure engineering platform that supports eCommerce, Digital Marketing, and IoT

Qualificações

  • 4+ years QA or ML testing engineering.
  • Experience with Python-based test automation using pytest.
  • Experience validating evaluator correctness with known-good / known-bad sessions.
  • Experience with on-demand mode integration testing in CI/CD.
  • Experience validating online mode sampling accuracy.
  • Knowledge of non-AgentCore runtime feasibility assessment methodology and AI/LLM system quality testing.

Responsabilidades

  • Design, implement, and maintain automated test suites for AI evaluation frameworks and evaluation pipelines.
  • Develop Python-based test automation using pytest to validate evaluator behavior, quality scoring, and framework reliability.
  • Create and maintain known-good and known-bad test datasets, sessions, and workflows for evaluator correctness validation.
  • Validate the accuracy and consistency of evaluation results across different agent workflows, prompts, tools, and execution scenarios.
  • Design and execute integration tests for on-demand evaluation workflows integrated into CI/CD pipelines.
  • Verify online evaluation behavior, sampling accuracy, and evaluation result consistency in production-like environments.
  • Collaborate with AI and platform teams to identify edge cases, failure scenarios, and evaluation blind spots.
  • Support feasibility assessments for applying evaluation frameworks to non-AgentCore runtimes and alternative AI execution environments.

Conhecimentos

QA testing
ML testing
Python
pytest
Evaluation framework testing
Automated testing

Ferramentas

CI/CD pipelines
OpenTelemetry
AWS

Descrição da oferta de emprego

We are looking for a QA / ML Tester – Evaluation Framework to ensure the quality, reliability, and correctness of evaluation systems used across enterprise AI and agent-based platforms. In this role, you will design and execute validation strategies, test evaluator behavior across a wide range of scenarios, and verify the accuracy of automated quality assessment pipelines. You will work closely with AI engineers, platform teams, and quality specialists to build confidence in evaluation results and support enterprise‑grade AI governance.

Our customer is a multinational corporation with more than a century of history and operations in over 180 countries. As part of its transformation, the company is driving the adoption of Reduced‑Risk Products (RRPs) for more than one billion consumers worldwide. Its technology landscape supports over 700 applications, requiring a scalable, secure, and highly reliable engineering platform.

At Intellias, you will join the Core Architecture Team, helping to engineer a comprehensive software ecosystem that powers best‑in‑class eCommerce, Digital Marketing, and IoT solutions. The Enterprise Platform provides engineering teams with shared services, technologies, and best practices that accelerate software delivery while ensuring quality, governance, compliance, and operational excellence across the software development lifecycle.

Requirements:
  • 4+ years QA or ML testing engineering
  • Python test automation (pytest)
  • Evaluator correctness testing (known-good/known-bad session pairs)
  • On-demand mode integration testing with CI/CD
  • Online mode sampling accuracy validation
  • Non-AgentCore runtime feasibility assessment methodology
  • AI/LLM system quality testing
Nice-to-have
  • AWS AgentCore Evaluation API testing
  • OpenTelemetry trace‑based evaluation input testing
  • Multi-evaluator execution correctness testing
Responsibilities:
  • Design, implement, and maintain automated test suites for AI evaluation frameworks and evaluation pipelines.
  • Develop Python-based test automation using pytest to validate evaluator behavior, quality scoring, and framework reliability.
  • Create and maintain known-good and known-bad test datasets, sessions, and workflows for evaluator correctness validation.
  • Validate the accuracy and consistency of evaluation results across different agent workflows, prompts, tools, and execution scenarios.
  • Design and execute integration tests for on-demand evaluation workflows integrated into CI/CD pipelines.
  • Verify online evaluation behavior, sampling accuracy, and evaluation result consistency in production‑like environments.
  • Conduct functional testing of evaluation components, including evaluator execution flows, scoring logic, and result aggregation.
  • Collaborate with AI and platform engineering teams to identify edge cases, failure scenarios, and evaluation blind spots.
  • Validate workflow compliance, tool execution assessment, and end‑to‑end quality evaluation processes.
  • Support feasibility assessments for applying evaluation frameworks to non‑AgentCore runtimes and alternative AI execution environments.
  • Analyze defects, inconsistencies, and quality regressions within evaluation systems and provide actionable recommendations.
  • Contribute to quality assurance standards, testing methodologies, and best practices for AI evaluation platforms.
Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Senior QA Engineer - AI Evaluation & Validation
Senior QA Engineer - AI Evaluation & Validation

Intellias • Portugal

Presencial
EUR 45 000 - 65 000
Platform Engineer - CI/CD Gate & Online Evaluation
Platform Engineer - CI/CD Gate & Online Evaluation

Intellias • Portugal

Presencial
EUR 60 000 - 90 000
Python Engineer - Evaluator Library
Python Engineer - Evaluator Library

Intellias • Portugal

Presencial
EUR 55 000 - 75 000
Annual leave 22 days
Sick leave 3 days
Permanent contract
+2
Python Engineer - AI Evaluator Library
Python Engineer - AI Evaluator Library

Intellias • Portugal

Presencial
EUR 55 000 - 75 000
Annual leave 22 days
Sick leave 3 days
Permanent contract
+2
Senior QA AI Engineer (Web Platforms & LLM Validation)
Senior QA AI Engineer (Web Platforms & LLM Validation)

Decskill • Lisboa

Híbrido
EUR 45 000 - 65 000
Senior Test Automation Engineer
Senior Test Automation Engineer

IgniteTech • Portugal

Presencial
EUR 55 000 - 75 000
Senior QA GEN AI Engineer
Senior QA GEN AI Engineer

99X Technology • Lisboa

Híbrido
EUR 45 000 - 65 000
AI QA Automation Engineer
AI QA Automation Engineer

Granter • Lisboa

Híbrido
EUR 35 000 - 55 000
QA Automation
QA Automation

Extia • Porto

Híbrido
EUR 35 000 - 60 000
Senior QA Engineer - AI/LLM Validation for Web Platforms
Senior QA Engineer - AI/LLM Validation for Web Platforms

Decskill • Lisboa

Híbrido
EUR 45 000 - 65 000