Python Engineer - Evaluator Library

Intellias

Portugal

Presencial

EUR 55 000 - 75 000

Tempo integral

Há 13 dias

Recebe mais respostas dos empregadores

Envia um currículo específico para a oferta em poucos minutos.

Vantagens oferecidas por esta oferta de emprego

Annual leave 22 days
Sick leave 3 days
Permanent contract
Health insurance
Office Porto

Resumo da oferta

Intellias is seeking a Python Engineer - Evaluator Library in Porto to design and implement reusable evaluation components for enterprise AI agents and LLM-powered workflows. You will build custom evaluation capabilities focused on automated quality validation, workflow compliance, PII protection, and structured output verification, collaborating with AI Platform, ML, and DevOps teams.

The role emphasizes creating deterministic quality checks, integrating with AWS Lambda and Bedrock, and

Qualificações

  • 4+ years Python engineering.
  • LLM evaluation or quality assurance for AI/ML systems.
  • AWS Lambda function development and deployment.

Responsabilidades

  • Design, develop, and maintain reusable Python-based evaluator libraries for AI agents and LLM-powered workflows.
  • Implement AWS Lambda-based custom evaluators to perform deterministic quality, compliance, and validation checks.
  • Develop LLM-as-a-judge evaluation logic to assess subjective dimensions such as relevance, helpfulness, consistency, and response quality.
  • Build automated PII detection evaluators using regex-based techniques and AWS Bedrock Guardrails integrations.
  • Implement TOOL_CALL-level validation mechanisms to verify structured outputs, JSON schema compliance, and tool response correctness.
  • Develop SESSION-level evaluators to validate workflow contract compliance, execution integrity, and cross-step behavioral expectations.
  • Create TRACE-level evaluators for numerical accuracy verification, calculation consistency, and deterministic result validation.
  • Integrate evaluator components with AWS AgentCore Evaluation workflows and enterprise AI quality pipelines.
  • Collaborate with AI platform teams to define evaluation standards, scoring methodologies, and quality acceptance criteria.
  • Design and maintain unit, integration, and validation tests for evaluator libraries and quality frameworks.
  • Support observability and troubleshooting by integrating evaluation outputs with CloudWatch logging and monitoring capabilities.

Conhecimentos

Python
LLM evaluation
AWS Lambda
PII detection
Quality assurance
Regex
AWS Bedrock
CloudWatch
Evaluation standards

Ferramentas

AWS Lambda
AWS Bedrock
Regex
CloudWatch

Descrição da oferta de emprego

Python Engineer - Evaluator Library to design and implement reusable evaluation components that ensure the quality, safety, and compliance of enterprise AI agents and LLM-powered workflows. You will build custom evaluation capabilities used across AI platforms, focusing on automated quality validation, workflow compliance, PII protection, and structured output verification. Working closely with AI Platform Engineers, ML Engineers, and DevOps teams, you will help establish reliable evaluation standards and scalable quality assurance mechanisms for agent-based systems.


Project Overview:



  • Our customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications.

  • Intellia's mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting-edge technology. Our teams are involved in the engineering of core platform components for best-in-class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform.

  • The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices.


Requirements:



  • Python (Lambda functions as AWS AgentCore custom code-based evaluators);

  • PII detection (regex-based + AWS Bedrock Guardrails);

  • Workflow contract compliance checking (SESSION level evaluator);

  • Numerical accuracy validation logic (TRACE level evaluator).


- Experience:



  • 4+ years Python engineering;

  • LLM evaluation or quality assurance for AI/ML systems;

  • AWS Lambda function development and deployment;


- Nice-to-have



  • AWS AgentCore Evaluation custom evaluator Lambda registration;

  • CloudWatch Logs as evaluator output sink.


Responsibilities:



  • Design, develop, and maintain reusable Python-based evaluator libraries for AI agents and LLM-powered workflows.

  • Implement AWS Lambda-based custom evaluators to perform deterministic quality, compliance, and validation checks.

  • Develop LLM-as-a-judge evaluation logic to assess subjective dimensions such as relevance, helpfulness, consistency, and response quality.

  • Build automated PII detection evaluators using regex-based techniques and AWS Bedrock Guardrails integrations.

  • Implement TOOL_CALL-level validation mechanisms to verify structured outputs, JSON schema compliance, and tool response correctness.

  • Develop SESSION-level evaluators to validate workflow contract compliance, execution integrity, and cross-step behavioral expectations.

  • Create TRACE-level evaluators for numerical accuracy verification, calculation consistency, and deterministic result validation.

  • Integrate evaluator components with AWS AgentCore Evaluation workflows and enterprise AI quality pipelines.

  • Collaborate with AI platform teams to define evaluation standards, scoring methodologies, and quality acceptance criteria.

  • Design and maintain unit, integration, and validation tests for evaluator libraries and quality frameworks.

  • Support observability and troubleshooting by integrating evaluation outputs with CloudWatch logging and monitoring capabilities.

  • Contribute to enterprise AI governance initiatives by improving evaluation coverage, auditability, and compliance controls.


What we offer:



  • 22 days of annual leave per year + 3 paid sick leaves;

  • Permanent contract;

  • Health Insurance (employee);

  • Modern, well-equipped office in the centre of Porto;

Obtém a tua avaliação gratuita e confidencial do currículo.
ou arrasta e larga o ficheiro aqui.
Similar jobs

Ofertas semelhantes que vale a pena comparar

Platform Engineer - CI/CD Gate & Online Evaluation
Platform Engineer - CI/CD Gate & Online Evaluation

Intellias • Portugal

Presencial
EUR 60 000 - 90 000
Python Engineer - AI Evaluator Library
Python Engineer - AI Evaluator Library

Intellias • Portugal

Presencial
EUR 55 000 - 75 000
Annual leave 22 days
Sick leave 3 days
Permanent contract
+2
Quality Assurance Tester
Quality Assurance Tester

Intellias • Portugal

Presencial
EUR 45 000 - 65 000
AI Platform Lead
AI Platform Lead

login.works • Alvalade

Híbrido
EUR 70 000 - 90 000
Direct access to leadership
Competitive compensation
Flexible working conditions
+1
Senior LLM Ops & AI Platform Engineer
Senior LLM Ops & AI Platform Engineer

Fyld • Portugal

Presencial
EUR 90 000 - 120 000
Lead AI Engineer
Lead AI Engineer

Indicium AI • Lisboa

Presencial
EUR 50 000 - 70 000
£2k annual training budget with 5 study days
Collaborative working environment
Financial support for new initiatives
+1
AI engineer
AI engineer

login.works • Alvalade

Híbrido
EUR 40 000 - 60 000
Competitive compensation
Flexible working hours
Mentorship opportunities
AI Engineer Lisbon
AI Engineer Lisbon

Indicium Tech • Lisboa

Presencial
EUR 50 000 - 70 000
Competitive salary
Company bonus
Personal learning budget
+2
AI Engineer
AI Engineer

ITSector • Castelo Branco

Híbrido
EUR 30 000 - 50 000
Health Insurance
Hybrid working
Technical training
+1
platform engineer for enterprise AI platforms
platform engineer for enterprise AI platforms

Enfint • Lisboa

Presencial
EUR 90 000 - 130 000