Senior Ai Agent Engineer

Firstignite

Arbo

Presencial

EUR 90.000 - 130.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura completa en un minuto: currículum y carta de presentación adaptados, listos para enviar.

Supera los filtros ATS

Descripción de la vacante

FirstIgnite busca un/a Senior AI Agent Engineer para construir agentes de IA en nuestro producto y desarrollar evaluaciones que midan mejoras. El puesto se centra en precisión y trazabilidad, con entregas que deben ser revisadas por un humano antes de cualquier acción.

Trabajarás con el Head of Engineering y equipos de producto. Se valorará experiencia previa en agentes, integración de APIs y manejo de pipelines de documentos complejos de manera eficiente.

Formación

  • Experiencia en ingeniería de modelos y agentes de IA que interactúan con usuarios.
  • Capacidad para evaluar agentes, no solo modelos, y entender métricas de run multistep.
  • Integración con APIs externas, incluso si la documentación es limitada.
  • Pipelines de documentos: extracción, normalización y verificación con fuente de verdad.
  • Diseño de experimentos para respuestas replicables y medibles.
  • Redacción clara y convincente para que las evaluaciones sean seguidas.

Responsabilidades

  • Diseñar y desplegar agentes de IA largos y multi‑pasos en SDKs y herramientas relevantes.
  • Crear y mantener herramientas para convertir APIs en llamadas seguras para agentes.
  • Extraer datos estructurados de documentos y compararlos con fuentes existentes.
  • Establecer suites de evaluación y medir exactitud de uso de herramientas y calidad de trayectoria.
  • Producir borradores con citaciones y señales de confianza para revisión humana.
  • Colaborar con producto y expertos para definir metas medibles a partir de datos limitados.
  • Transformar feedback en mejoras de rendimiento y evaluaciones reproducibles.
  • Escribir templates y documentación para que el equipo ejecute evaluaciones.

Conocimientos

LLM engineering experience
Agent evaluation
API integration
Document pipelines
Experiment design
Clear written communication

Herramientas

Temporal
LangGraph
OpenAI Agents SDK
Braintrust

Descripción del empleo

Descripción del trabajo About FirstIgnite FirstIgnite makes software for university tech transfer offices. Those are the people who take research coming out of a university lab and get it patented, licensed, or spun out into a company. ¿Posee las habilidades y la experiencia adecuadas para este puesto? Siga leyendo para descubrirlo y envíe su solicitud. The role We’re hiring a Senior AI Agent Engineer. You’ll build the agents in our product, and you’ll build the evals that tell us whether each change made them better or worse. The work is document-heavy rather than chat. The agents run multi-step, call tools, read long and inconsistently formatted source material, check it against existing records, and produce output that a person reviews before anything happens with it. Accuracy matters more here than speed or novelty. Most of the engineering effort goes into precision, traceability, and getting the agent to hand off to a human at the right moment. You’ll report to the Head of Engineering and work with product and the full-stack team. If you’ve shipped agents before, you’ve probably had the experience of changing a prompt and having no idea whether you improved anything. That problem is most of this job.

What you’ll do
  • Design and ship long-running, multi-step, tool-using agents on various AI SDKs and tooling, included but not limited to the OpenAI Agents SDK, the Anthropic Agent SDK, the Vercel AI SDK, LangGraph, MCP, and Temporal Cloud.
  • Wrap our APIs and our partners’ APIs as tools an agent can call over MCP. Some of those systems are old, single-tenant, and outside our control, so a fair amount of the work is translation.
  • Get structured data out of long documents and match it against records that already exist. Expect entity resolution and fuzzy matching, and expect much of it to run in batch.
  • Stand up eval suites using various evaluation frameworks and tooling, included but not limited to Promptfoo, Braintrust, LangSmith, DeepEval, LLM-as-judge methods, and custom harnesses. Measure tool-use correctness, trajectory quality, and whether the agent finished the task.
  • Every agent here produces a draft that a person signs off on. Build the citations and confidence signals that make that review fast, and give the agent a clear way to escape.
  • Sit with product and domain experts and turn vague quality goals into something measurable. Sometimes the only dataset available for that is tiny, or confidential, or both.
  • Instrument production traffic, turn real customer interactions into golden datasets, and run them as regression tests.
  • Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt strategies and agent designs, and know what each option costs in latency and quality.
  • Bootstrap quality signal for features that have no production traffic yet. That usually means generating synthetic documents and test cases, including the ugly edge cases real customers will eventually send us, and knowing where synthetic data stops being a good proxy.
  • Write the templates, docs, and tooling the rest of the team needs to run evals without coming to you.
Required Qualifications
  • 3+ years of engineering experience, including hands‑on work on LLM or agent systems that real users touched.
  • You’ve evaluated agents, not only models, and you know why single‑turn accuracy says little about a multi‑step run.
  • You’ve integrated against APIs you don’t own, including old ones with bad documentation, and turned them into something an agent can call reliably.
  • You’re comfortable with document pipelines: pulling data out, normalizing it, and checking it against a structured source of truth.
  • You’ve used at least one LLM evaluation framework, in‑house tooling included.
  • You know how LLM-as‑judge methods break down (position bias, verbosity bias, judge drift) and what to do about it.
  • You can tell a real regression from noise, and design an experiment that answers the question being asked instead of a nearby one.
  • You can read a customer call transcript, work out which failures matter, and ship a fix and an eval for them.
  • You write clearly. Engineers won’t act on eval results they don’t read or don’t trust.
  • You’re based somewhere between New York time (ET) and Western European time. Italy is the furthest east we can go.
  • You might currently be titled AI Agent Engineer · Agentic AI Engineer · Senior AI Engineer · AI Systems Engineer · Lead AI Developer · AI Solutions Architect · AI Integration Specialist · Senior Data Scientist, AI Titles are all over the place in this space. If the work above matches what you already do, apply. We’ll go on what you’ve shipped.
Preferred Qualifications
  • You’ve evaluated retrieval systems: RAG, hybrid search, reranking.
  • You’ve worked with agent orchestration frameworks like Temporal, LangGraph, or the OpenAI Agents SDK, and you know how long‑running tool use goes wrong.
  • You have a background in information retrieval or search relevance.
  • You’ve worked somewhere an agent’s output carried financial or compliance consequences.
  • You’ve built internal tooling that non‑engineers used on their own to label and review model output.

xiphteb This is a fully remote, full-time permanent position available to candidates located within the New York (ET) through Western Europe time zones, with flexible working hours to support collaboration across regions.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior AI Agent Engineer
Senior AI Agent Engineer

FirstIgnite • O Chazo

Presencial
EUR 90.000 - 130.000
Senior AI Agent Engineer
Senior AI Agent Engineer

FirstIgnite • As Catro Rúas

Presencial
EUR 70.000 - 95.000
Senior AI Agent Engineer
Senior AI Agent Engineer

FirstIgnite • Valencia

Presencial
EUR 60.000 - 90.000
Senior AI Agent Engineer
Senior AI Agent Engineer

FirstIgnite • Vitoria

Presencial
EUR 70.000 - 90.000
Generative Ai Engineer: Agent-Driven Production
Generative Ai Engineer: Agent-Driven Production

Wizeline • Salamanca

Presencial
EUR 60.000 - 90.000
AI Software Engineer | Spain
AI Software Engineer | Spain

Accenture España • Madrid

Presencial
EUR 90.000 - 130.000
Senior AI Agent Engineer - Remote, Precision-Focused
Senior AI Agent Engineer - Remote, Precision-Focused

FirstIgnite • España

Presencial
EUR 90.000 - 130.000
AI Software Engineer
AI Software Engineer

Accenture España • Madrid

Presencial
EUR 65.000 - 95.000
Travel opportunities
AI Engineer (Full-Stack)
AI Engineer (Full-Stack)

Bifrost Studios • Barcelona

Presencial
Confidential
Security and GDPR compliance
Supabase in production
Multi-tenant SaaS experience
+1
Senior AI Agent Engineer - Precision, Tooling & Evaluation
Senior AI Agent Engineer - Precision, Tooling & Evaluation

FirstIgnite • O Chazo

Presencial
EUR 90.000 - 130.000