Senior AI Quality & Evaluation Engineer — Remote

Intellias

Ibiza

Presencial

EUR 90.000 - 130.000

Jornada completa

Hace 5 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo: un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Descripción de la vacante

Intellias seeks a Senior ML / Evaluation Engineer to define and implement enterprise-grade evaluation standards for AI agents and LLM-powered applications. You will design evaluation frameworks, build custom evaluation pipelines, and establish automated quality gates across the AI delivery lifecycle.

Collaborate with AI Platform Engineers, ML Engineers, and DevOps to ensure reliable, measurable, and production-ready AI systems through scalable evaluation, observability, and governance practices.

Formación

  • 5+ years ML engineering or AI platform engineering
  • LLM evaluation framework design and implementation
  • Custom evaluator implementation for deterministic quality checks
  • CI/CD deployment gate design for ML model or agent quality
  • AWS AgentCore Evaluation (on-demand mode for CI/CD gates, online mode for production sampling)
  • LLM-as-judge evaluator design (built-in AgentCore evaluators — helpfulness, correctness)
  • Custom code-based Lambda evaluators (Python — deterministic checks)
  • Evaluation levels (TRACE for per-response, TOOL CALL for per-invocation, SESSION for workflow)
  • OTel spans from AWS AgentCore Observability as evaluation input
  • Enterprise evaluation standard authoring (mandatory dimensions, pass/fail criteria)

Responsabilidades

  • Design, implement, and maintain enterprise-grade evaluation frameworks for LLMs, AI agents, and multi-step AI workflows.
  • Develop and optimize LLM-as-a-judge evaluators to assess dimensions such as helpfulness, correctness, consistency, and policy compliance.
  • Build custom Python-based evaluators using AWS Lambda to perform deterministic validation, business-rule enforcement, and workflow quality checks.
  • Define and implement evaluation standards, mandatory quality dimensions, scoring methodologies, and pass/fail criteria across AI platforms.
  • Design evaluation strategies at multiple levels, including TRACE, TOOL CALL, and SESSION evaluation scopes.
  • Integrate evaluation workflows into CI/CD pipelines and establish automated deployment quality gates for AI-powered applications.
  • Leverage AWS AgentCore Evaluation capabilities to execute on-demand evaluations and support production quality monitoring.
  • Utilize observability data, OpenTelemetry traces, and AgentCore telemetry signals as evaluation inputs for quality assessment and root-cause analysis.
  • Collaborate with platform, security, and AI engineering teams to improve agent reliability, accuracy, and operational quality.
  • Analyze evaluation results, identify quality regressions, and drive corrective actions across models, prompts, tools, and workflows.
  • Define monitoring and reporting mechanisms for evaluation outcomes, quality trends, and operational KPIs.
  • Contribute to the evolution of enterprise AI governance, testing methodologies, and evaluation best practices.

Conocimientos

ML/AI platform engineering
LLM evaluation
Custom evaluators
CI/CD gates
AWS AgentCore
LLM-as-judge evaluators
Python Lambda evaluators
Evaluation levels
Observability inputs
Evaluation standards

Herramientas

AWS AgentCore
OpenTelemetry
Python

Descripción del empleo

Intellias seeks a Senior ML / Evaluation Engineer to define and implement enterprise-grade evaluation standards for AI agents and LLM-powered applications. You will design evaluation frameworks, build custom evaluation pipelines, and establish automated quality gates across the AI delivery lifecycle.

Collaborate with AI Platform Engineers, ML Engineers, and DevOps to ensure reliable, measurable, and production-ready AI systems through scalable evaluation, observability, and governance practices.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior ML Evaluation Engineer — Enterprise AI Quality Remote
Senior ML Evaluation Engineer — Enterprise AI Quality Remote

Intellias • Almería

Presencial
EUR 70.000 - 110.000
Senior ML Evaluation Engineer — Remote (Spain)
Senior ML Evaluation Engineer — Remote (Spain)

Intellias • Alicante

Presencial
EUR 70.000 - 100.000
Senior ML Evaluation Engineer - Enterprise AI Quality Remote
Senior ML Evaluation Engineer - Enterprise AI Quality Remote

Intellias • País Vasco

Presencial
EUR 75.000 - 110.000
Senior ML Evaluation Engineer — Remote (Spain)
Senior ML Evaluation Engineer — Remote (Spain)

Intellias • Córdoba

Presencial
EUR 90.000 - 130.000
Remote Platform Engineer: AI Quality Gates & CI/CD
Remote Platform Engineer: AI Quality Gates & CI/CD

Intellias • Huelva

Presencial
EUR 70.000 - 90.000
Platform Engineer — Ci/Cd Gate & Online Evaluation - Intellias
Platform Engineer — Ci/Cd Gate & Online Evaluation - Intellias

Intellias • Navarra

Presencial
EUR 60.000 - 90.000
AI Platform Quality Engineer - CI/CD Gates & Evaluation
AI Platform Quality Engineer - CI/CD Gates & Evaluation

Intellias • Gijón

Presencial
EUR 55.000 - 90.000
Senior ML Evaluation Engineer - AI Quality Gates
Senior ML Evaluation Engineer - AI Quality Gates

Intellias • Barcelona

Presencial
EUR 90.000 - 130.000
Platform Engineer: CI/CD Gate & AI Quality Automation
Platform Engineer: CI/CD Gate & AI Quality Automation

Intellias • Badajoz

Presencial
EUR 70.000 - 100.000
Senior AI Engineer: Scalable, Production-Ready AI (Remote)
Senior AI Engineer: Scalable, Production-Ready AI (Remote)

UL Solutions • Barcelona

Híbrido
EUR 53.000 - 58.000