Senior ML / Evaluation Engineer

Intellias

Spain (TX)

On-site

EUR 70,000 - 100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Intellias is seeking a Senior ML / Evaluation Engineer to help define and implement quality standards for enterprise-grade AI agents and LLM-powered applications. You will design evaluation frameworks, build custom evaluation pipelines, and establish automated quality gates across the AI delivery lifecycle.

You will work with AI Platform Engineers, ML Engineers, and DevOps teams to ensure production-ready AI systems.

Qualifications

  • 5+ years ML engineering or AI platform engineering.
  • LLM evaluation framework design and implementation.
  • Custom evaluator implementation for deterministic quality checks.
  • AWS AgentCore Evaluation (on-demand mode for CI/CD gates, online mode for production sampling).
  • Custom code-based Lambda evaluators (Python - deterministic checks).
  • Evaluation levels (TRACE, TOOL_CALL, SESSION).
  • OTel spans from AWS AgentCore Observability as evaluation input.

Responsibilities

  • Design, implement, and maintain enterprise-grade evaluation frameworks for LLMs and AI agents.
  • Build and optimize LLM-as-a-judge evaluators for multiple dimensions (helpfulness, correctness, consistency, policy).
  • Develop Python-based evaluators using AWS Lambda for deterministic validation and workflow checks.
  • Define evaluation standards, scoring methodologies, and pass/fail criteria across AI platforms.
  • Design evaluation strategies at TRACE, TOOL_CALL, and SESSION levels.
  • Integrate evaluation workflows into CI/CD pipelines with automated quality gates.
  • Leverage AWS AgentCore and OpenTelemetry signals as evaluation inputs for quality analysis.
  • Collaborate with platform, security, and AI engineering teams to improve reliability.

Skills

ML engineering
AI platform engineering
LLM evaluation
Custom evaluators
AWS AgentCore
Python evaluators
Evaluation levels
Observability data

Tools

Python
AWS Lambda
OpenTelemetry
AgentCore
CI/CD tooling

Job description

Location: Remote from Spain (an indefinite Spanish employment contract)

Senior ML / Evaluation Engineer to help define and implement quality standards for enterprise-grade AI agents and LLM-powered applications. In this role, you will design evaluation frameworks, build custom evaluation pipelines, and establish automated quality gates across the AI delivery lifecycle. You will work closely with AI Platform Engineers, ML Engineers, and DevOps teams to ensure reliable, measurable, and production-ready AI systems through scalable evaluation, observability, and governance practices.

Project Overview

Our customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications.

Intellia's mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting-edge technology. Our teams are involved in the engineering of core platform components for best-in-class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform.

The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices.

Requirements
  • 5+ years ML engineering or AI platform engineering
  • LLM evaluation framework design and implementation
  • Custom evaluator implementation for deterministic quality checks
  • AWS AgentCore Evaluation (on-demand mode for CI/CD gates, online mode for production sampling)
  • Custom code-based Lambda evaluators (Python - deterministic checks)
  • Evaluation levels (TRACE for per-response, TOOL_CALL for per-invocation, SESSION for workflow)
  • OTel spans from AWS AgentCore Observability as evaluation input
Nice-to-have
  • AWS Bedrock Guardrails for PII detection evaluator integration
  • CloudWatch metrics output from AgentCore Evaluation for online mode
Responsibilities
  • Design, implement, and maintain enterprise-grade evaluation frameworks for LLMs, AI agents, and multi-step AI workflows.
  • Develop and optimize LLM-as-a-judge evaluators to assess dimensions such as helpfulness, correctness, consistency, and policy compliance.
  • Build custom Python-based evaluators using AWS Lambda to perform deterministic validation, business-rule enforcement, and workflow quality checks.
  • Define and implement evaluation standards, mandatory quality dimensions, scoring methodologies, and pass/fail criteria across AI platforms.
  • Design evaluation strategies at multiple levels, including TRACE, TOOL_CALL, and SESSION evaluation scopes.
  • Integrate evaluation workflows into CI/CD pipelines and establish automated deployment quality gates for AI-powered applications.
  • Leverage AWS AgentCore Evaluation capabilities to execute on-demand evaluations and support production quality monitoring.
  • Utilize observability data, OpenTelemetry traces, and AgentCore telemetry signals as evaluation inputs for quality assessment and root-cause analysis.
  • Collaborate with platform, security, and AI engineering teams to improve agent reliability, accuracy, and operational quality.
  • Analyze evaluation results, identify quality regressions, and drive corrective actions across models, prompts, tools, and workflows.
  • Define monitoring and reporting mechanisms for evaluation outcomes, quality trends, and operational KPIs.
  • Contribute to the evolution of enterprise AI governance, testing methodologies, and evaluation best practices.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Town of Poland (NY)

On-site
USD 140,000 - 200,000
Python Engineer — Evaluator Library
Python Engineer — Evaluator Library

Intellias • Town of Poland (NY)

On-site
USD 120,000 - 180,000
QA / ML Tester — Evaluation Framework
QA / ML Tester — Evaluation Framework

Intellias • Town of Poland (NY)

On-site
USD 120,000 - 160,000
Remote Senior ML Evaluation Engineer – AI Quality
Remote Senior ML Evaluation Engineer – AI Quality

Intellias • Spain (TX)

On-site
EUR 70,000 - 100,000
Senior ML Evaluation Engineer - LLM Quality & Gates
Senior ML Evaluation Engineer - LLM Quality & Gates

Intellias • Town of Poland (NY)

On-site
USD 140,000 - 200,000
Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Evaluation Engineer
AI Evaluation Engineer

Capital Rx • Denver (CO)

On-site
USD 120,000 - 180,000
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
AI Evaluation Engineer
AI Evaluation Engineer

Capital Rx • Charlotte (NC)

On-site
USD 120,000 - 180,000