Python Engineer — Evaluator Library

Intellias

Town of Poland (NY)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Intellias seeks a Python Engineer — Evaluator Library to design and implement reusable evaluation components that ensure quality, safety, and compliance of enterprise AI agents and LLM-powered workflows. You will build capabilities used across AI platforms, focusing on automated quality validation and structured output verification.

As part of the Core Architecture Team, you will develop Lambda-based evaluators, PII detection, and governance controls, collaborating with AI Platform, ML, and

Qualifications

  • 4+ years Python engineering experience.
  • LLM evaluation or AI/ML QA experience.
  • AWS Lambda development and deployment experience.

Responsibilities

  • Design, develop, and maintain reusable Python-based evaluator libraries for AI agents.
  • Implement AWS Lambda-based evaluators to perform quality, compliance, and validation checks.
  • Develop LLM-as-a-judge evaluation logic for relevance, helpfulness, and consistency.
  • Build PII detection evaluators using regex and Bedrock Guardrails integrations.
  • Implement SESSION- and TRACE-level evaluators for workflow contracts and numerical accuracy.
  • Integrate evaluator components with AWS AgentCore Evaluation workflows and enterprise QA pipelines.

Skills

Python
PII detection
Workflow contract compliance checking
Numerical accuracy validation
AWS Lambda

Tools

AWS Bedrock
CloudWatch
Regex

Job description

We are looking for a Python Engineer — Evaluator Library to design and implement reusable evaluation components that ensure the quality, safety, and compliance of enterprise AI agents and LLM-powered workflows. You will build custom evaluation capabilities used across AI platforms, focusing on automated quality validation, workflow compliance, PII protection, and structured output verification. Working closely with AI Platform Engineers, ML Engineers, and DevOps teams, you will help establish reliable evaluation standards and scalable quality assurance mechanisms for agent-based systems.

Project Overview

Our customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications.

Intellia's mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting-edge technology. Our teams are involved in the engineering of core platform components for best-in-class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform.

The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices.

Requirements
Skills
  • Python (Lambda functions as AWS AgentCore custom code-based evaluators)
  • PII detection (regex-based + AWS Bedrock Guardrails)
  • Workflow contract compliance checking (SESSION level evaluator)
  • Numerical accuracy validation logic (TRACE level evaluator)
Experience
  • 4+ years Python engineering
  • LLM evaluation or quality assurance for AI/ML systems
  • AWS Lambda function development and deployment
Nice-to-have
  • AWS Bedrock Guardrails for PII detection integration
  • CloudWatch Logs as evaluator output sink
Responsibilities
  • Design, develop, and maintain reusable Python-based evaluator libraries for AI agents and LLM-powered workflows.
  • Implement AWS Lambda-based custom evaluators to perform deterministic quality, compliance, and validation checks.
  • Develop LLM-as-a-judge evaluation logic to assess subjective dimensions such as relevance, helpfulness, consistency, and response quality.
  • Build automated PII detection evaluators using regex-based techniques and AWS Bedrock Guardrails integrations.
  • Implement TOOL_CALL-level validation mechanisms to verify structured outputs, JSON schema compliance, and tool response correctness.
  • Develop SESSION-level evaluators to validate workflow contract compliance, execution integrity, and cross-step behavioral expectations.
  • Create TRACE-level evaluators for numerical accuracy verification, calculation consistency, and deterministic result validation.
  • Integrate evaluator components with AWS AgentCore Evaluation workflows and enterprise AI quality pipelines.
  • Collaborate with AI platform teams to define evaluation standards, scoring methodologies, and quality acceptance criteria.
  • Design and maintain unit, integration, and validation tests for evaluator libraries and quality frameworks.
  • Support observability and troubleshooting by integrating evaluation outputs with CloudWatch logging and monitoring capabilities.
  • Contribute to enterprise AI governance initiatives by improving evaluation coverage, auditability, and compliance controls.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Spain (TX)

On-site
EUR 70,000 - 100,000
Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Town of Poland (NY)

On-site
USD 140,000 - 200,000
QA / ML Tester — Evaluation Framework
QA / ML Tester — Evaluation Framework

Intellias • Town of Poland (NY)

On-site
USD 120,000 - 160,000
Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Evaluation Engineer
AI Evaluation Engineer

Capital Rx • Charlotte (NC)

On-site
USD 120,000 - 180,000
AI Evaluation Engineer
AI Evaluation Engineer

Capital Rx • Denver (CO)

On-site
USD 120,000 - 180,000
AI Evaluation Engineer
AI Evaluation Engineer

Capital Rx • New York (NY)

On-site
USD 120,000 - 160,000
Forward Deployed Engineer, Applied AI
Forward Deployed Engineer, Applied AI

Jobtailor • California (MO)

Hybrid
USD 180,000 - 270,000
AI Evaluation Scientist
AI Evaluation Scientist

Steampunk, Inc. • McLean (VA)

On-site
USD 140,000 - 210,000
Senior AI Engineer
Senior AI Engineer

arosplatforms | AI Consulting & Services • North Township (IN)

On-site
USD 120,000 - 170,000