Lead Engineer, AI Platform

Lever, Inc.

España

A distancia

EUR 129.000 - 174.000

Jornada completa

hace 35 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Remote work environment
35 days PTO per year
Equity in the company
Twice-yearly international retreats
Health and wellbeing benefits
Professional development
Global, distributed team

Descripción de la vacante

Lever, Inc. seeks a Lead Engineer, AI Platform based in Spain to lead the engineering foundation behind reliable, measurable AI-powered features. You will build evaluation frameworks, observability tooling, and diagnostic infrastructure for production AI agents.

You’ll combine hands-on software development with technical leadership, managing an AI Quality team and collaborating with AI Core teams to improve product quality.

Formación

  • 7+ years of experience building and shipping production software, ideally with LLM-powered agents.
  • Experience with complex tool-using AI systems with planning, orchestration, or sub-agents.
  • Strong ability to demonstrate shipped software and explain how effects were measured.
  • Experience with Ruby on Rails and/or Python; quick productivity in new tech.
  • Experience building evaluation or observability infrastructure for ML/AI systems.
  • Familiarity with evaluation frameworks such as Braintrust, LangSmith, or similar tools.
  • Experience designing datasets, annotation workflows, or labeling pipelines for ML/AI evaluation.
  • Ability to learn quickly, experiment extensively, and use empirical results to guide decisions.
  • Comfortable in fast-paced, ambiguous environments with evolving requirements.
  • Strong technical leadership and people-management capabilities.
  • Excellent English communication (CEFR C2 / ILR 5).
  • Strong alignment with a collaborative, ownership-oriented culture.

Responsabilidades

  • Design, build, and own evaluation infrastructure, incl. CI/CD pipelines, scorers, datasets, and multi-turn evaluation.
  • Develop observability and diagnostic capabilities across planning, execution, tool selection, and agent trajectories.
  • Investigate failures in AI workflows and turn findings into prototypes or prioritized improvements.
  • Build datasets via human annotation, AI-generated examples, and simulated conversations.
  • Develop experimentation frameworks for prompts, models, and agent harnesses against baselines.
  • Identify cost and latency optimization opportunities through model selection and caching.
  • Set technical direction and manage day-to-day priorities for the AI Quality team while staying hands-on.
  • Partner with AI Core teams to measure changes and prove quality improvements.
  • Establish engineering practices and evaluation approaches for reliable production AI systems.

Conocimientos

LLM-powered agents
Observability infrastructure
Evaluation pipelines
Ruby on Rails
Python
Leadership
People management
English proficiency

Herramientas

Braintrust
LangSmith
CI/CD tooling

Descripción del empleo

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Engineer, AI Platform based in Spain.

This role offers the opportunity to lead the engineering foundation behind reliable, measurable, and scalable AI-powered features. You’ll build evaluation frameworks, observability tooling, and diagnostic infrastructure that reveal how AI agents perform in real production environments. The position combines hands-on software engineering with technical leadership and people management. You’ll investigate quality issues across complex agent workflows, develop datasets and evaluation systems, and run experiments across models, prompts, and agent architectures. You’ll also help optimize AI systems for cost, latency, reliability, and overall user experience. Working in a highly remote and asynchronous environment, you’ll collaborate closely with AI engineering teams while shaping the technical direction of a growing AI Quality function.

Accountabilities
  • Design, build, and own evaluation infrastructure, including CI/CD pipelines, scorers, datasets, and systems for assessing AI agents from individual tool calls through complete multi-turn conversations.
  • Develop observability and diagnostic capabilities to identify exactly where quality issues occur across planning, execution, tool selection, and complex agent trajectories.
  • Investigate failures across sophisticated AI workflows and turn findings into technical prototypes, improvements, or clearly defined priorities for AI engineering teams.
  • Build and expand datasets through human annotation, AI-generated examples, and simulated conversations to increase evaluation coverage efficiently.
  • Develop structured experimentation frameworks for prompts, models, and agent harnesses, including evaluation of new and open-source models against production baselines.
  • Identify opportunities to improve AI system cost and latency through model selection, caching, routing, and other optimization strategies.
  • Set the technical direction and manage day-to-day priorities for the AI Quality engineering team while remaining actively involved in hands-on development.
  • Partner closely with AI Core engineering teams to ensure changes to AI products can be measured effectively and demonstrably improve quality.
  • Establish engineering practices and evaluation approaches that support reliable, efficient, and scalable production AI systems.
Requirements
  • 7+ years of experience building and shipping production software, ideally including LLM-powered agents capable of taking real actions within products.
  • Experience working with complex, tool-using AI systems involving multiple tools, planning, orchestration, or sub-agents rather than only simple, single-turn assistants.
  • Strong ability to demonstrate shipped software and explain how its effectiveness and reliability were measured.
  • Experience with Ruby on Rails and/or Python, with the ability to become productive quickly in technologies that may be new to you.
  • Experience building evaluation or observability infrastructure for ML/AI systems, including evaluation pipelines, scorers, dashboards, or CI/CD systems for evaluations.
  • Familiarity with evaluation frameworks such as Braintrust, LangSmith, or similar tools.
  • Experience designing datasets, annotation workflows, or labeling pipelines for machine learning or AI evaluation.
  • Ability to learn quickly, experiment extensively, and use empirical results to guide technical decisions.
  • Comfortable operating in a fast-paced environment with ambiguity and changing technical requirements.
  • Strong technical leadership and people-management capabilities, with the ability to balance team leadership and hands-on engineering.
  • Excellent English proficiency in spoken, written, and reading communication, equivalent to CEFR C2 / ILR 5.
  • Strong alignment with a collaborative, ownership-oriented engineering culture.
Benefits
  • Annual cash compensation of $170,000 USD, benchmarked to U.S. compensation levels regardless of location.
  • Equity in the company, including ongoing refresh grants.
  • 35 days of paid time off per year.
  • Fully remote work environment.
  • Significant flexibility and autonomy in how you organize your work.
  • Twice-yearly company retreats in international destinations.
  • Benefits supporting health, wellbeing, and professional development.
  • Opportunity to lead and grow an AI Quality engineering team while remaining hands-on technically.
  • Exposure to advanced AI agents, evaluation infrastructure, observability, experimentation, and production AI optimization.
  • Opportunity to work with a globally distributed team across multiple countries and time zones.
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Lead AI Engineer
Lead AI Engineer

InteractiveAI Limited • Madrid

Presencial
EUR 110.000 - 130.000
Equity plan
Health & wellness allowances
Private health insurance
+3
Expert Team Lead, Engineering
Expert Team Lead, Engineering

Lever, Inc. • España

Híbrido
EUR 70.000 - 120.000
Competitive compensation
Remote-friendly options
Lead Software Engineer - Working with AI
Lead Software Engineer - Working with AI

Harnham • España

Presencial
EUR 80.000 - 120.000
Competitive salary + equity/benefits package
Flexible working arrangements
Opportunity to work with cutting-edge technologies
Lead Data & AI Engineer - GenAI & AI Platforms
Lead Data & AI Engineer - GenAI & AI Platforms

Value Crew • Barcelona

Presencial
EUR 75.000 - 90.000
Hybrid work model
Private health insurance
Equity package
Lead QA Engineer (AI native)
Lead QA Engineer (AI native)

Lever, Inc. • España

A distancia
EUR 70.000 - 120.000
Fully remote
28 vacation days
Wellness days
+3
(Senior or Staff) Backend Engineer, AI tooling
(Senior or Staff) Backend Engineer, AI tooling

Lever, Inc. • España

Híbrido
EUR 117.000 - 143.000
RSUs
25 days annual leave
Visa sponsorship may be available for
AI Engineer
AI Engineer

Clarity • Madrid

Presencial
EUR 70.000 - 105.000
Competitive pay
Location flexibility
Generous time off
+4
Senior/Lead AI Engineer – Agents Team - BCN/MAD/MUN
Senior/Lead AI Engineer – Agents Team - BCN/MAD/MUN

Personio • Madrid

Híbrido
EUR 70.000 - 120.000
Hybrid work model
Remote AI Tooling Engineer - Developer Platforms
Remote AI Tooling Engineer - Developer Platforms

Jobgether • Madrid

Presencial
EUR 70.000 - 120.000
AI Engineer
AI Engineer

Trust In SODA • Madrid

Híbrido
EUR 90.000 - 130.000
Flexible hybrid working model
Private healthcare
Enhanced family leave policies
+1