Python AI Evaluator Engineer — Quality & Compliance

Intellias

Poland

On-site

PLN 180,000 - 240,000

Full time

11 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Intellias is seeking a Python Engineer — Evaluator Library to design and implement reusable evaluation components for enterprise AI agents and LLM-powered workflows. You will build custom evaluators focusing on automated quality validation, workflow compliance, PII protection, and structured output verification.

Collaborate with AI Platform Engineers, ML Engineers, and DevOps to establish evaluation standards and scalable quality assurance across agent-based systems.

Qualifications

  • 4+ years Python engineering experience.
  • LLM evaluation or quality assurance for AI/ML systems.
  • AWS Lambda function development and deployment.

Responsibilities

  • Design, develop, and maintain reusable Python-based evaluator libraries for AI agents and LLM-powered workflows.
  • Implement AWS Lambda-based custom evaluators to perform deterministic quality, compliance, and validation checks.
  • Develop LLM-as-a-judge evaluation logic to assess relevance, helpfulness, consistency, and response quality.
  • Build automated PII detection evaluators using regex-based techniques and AWS Bedrock Guardrails integrations.
  • Implement TOOL_CALL-level validation mechanisms to verify structured outputs, JSON schema compliance, and tool response correctness.
  • Develop SESSION-level evaluators to validate workflow contract compliance, execution integrity, and cross-step behavioral expectations.
  • Create TRACE-level evaluators for numerical accuracy verification, calculation consistency, and deterministic result validation.
  • Integrate evaluator components with AWS AgentCore Evaluation workflows and enterprise AI quality pipelines.
  • Collaborate with AI platform teams to define evaluation standards, scoring methodologies, and quality acceptance criteria.
  • Design and maintain unit, integration, and validation tests for evaluator libraries and quality frameworks.
  • Support observability and troubleshooting by integrating evaluation outputs with CloudWatch logging and monitoring capabilities.
  • Contribute to enterprise AI governance initiatives by improving evaluation coverage, auditability, and compliance controls.

Skills

Python
PII detection
Workflow contract compliance checking
Numerical accuracy validation

Tools

AWS Lambda
AWS Bedrock Guardrails
CloudWatch Logs

Job description

Intellias is seeking a Python Engineer — Evaluator Library to design and implement reusable evaluation components for enterprise AI agents and LLM-powered workflows. You will build custom evaluators focusing on automated quality validation, workflow compliance, PII protection, and structured output verification.

Collaborate with AI Platform Engineers, ML Engineers, and DevOps to establish evaluation standards and scalable quality assurance across agent-based systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Python Engineer — Evaluator Library
Python Engineer — Evaluator Library

Intellias • Poland

On-site
PLN 180,000 - 240,000
Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Poland

On-site
PLN 180,000 - 240,000
Senior ML Evaluation Architect for Enterprise AI
Senior ML Evaluation Architect for Enterprise AI

Intellias • Poland

On-site
PLN 180,000 - 240,000
Senior QA Engineer – ML/AI Evaluation Pipelines
Senior QA Engineer – ML/AI Evaluation Pipelines

EPAM Systems • Poland

Hybrid
PLN 180,000 - 260,000
Hybrid work design
Work remotely within Poland
Opportunity to work abroad up to 60/90
Python AI Agent Engineer for Enterprise Automation
Python AI Agent Engineer for Enterprise Automation

Intetics • Poland

Hybrid
PLN 180,000 - 240,000
1122 | Python/AI Agent Engineer
1122 | Python/AI Agent Engineer

Intetics • Poland

Hybrid
PLN 180,000 - 240,000
Senior AI QA Engineer - Automation & ML Evaluation
Senior AI QA Engineer - Automation & ML Evaluation

Sii Poland • Poznań

On-site
PLN 180,000 - 280,000
1122 | Python/AI Agent Engineer
1122 | Python/AI Agent Engineer

Intetics • Warszawa

On-site
PLN 180,000 - 240,000
Senior AI Platform Engineer - Scalable LLM Ops
Senior AI Platform Engineer - Scalable LLM Ops

EPAM Systems • Poland

Hybrid
PLN 180,000 - 260,000
Hybrid by design
Remote work within Poland
Relocation opportunities
+7
AI / LLM Engineer
AI / LLM Engineer

Link Group • Poland

On-site
PLN 180,000 - 240,000