Senior AI Agent Engineer

FirstIgnite

As Catro Rúas

Presencial

EUR 70.000 - 95.000

Jornada completa

Hace 9 días
Generador de candidaturas

Transforma esta oferta en una entrevista: un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Descripción de la vacante

FirstIgnite is hiring a Senior AI Agent Engineer to build multi-step, tool-using agents that work with long, document-heavy inputs. You will design, implement, and evaluate agent workflows, wrapping external APIs and ensuring precise, auditable outputs that can be reviewed before any action is taken.

You will work with product and the full-stack team to tighten evaluation, improve accuracy, and hand off to humans at the right moment.

Formación

  • 3+ years engineering experience with LLM or agent systems touched by real users.
  • Evaluated agents, not just models, and understand multi-step runs beyond single-turn accuracy.
  • Integrated against external APIs and translated them into reliable agent-callable interfaces.
  • Comfortable with document pipelines: extraction, normalization, and matching to truth sources.
  • Familiar with at least one LLM evaluation framework and in-house tooling.

Responsabilidades

  • Design and ship long-running, multi-step, tool-using agents across multiple AI SDKs.
  • Wrap APIs and partner APIs as tools a agent can call over MCP, including translation when needed.
  • Extract structured data from long documents and match against existing records.
  • Set up eval suites and measure tool-use correctness and trajectory quality.
  • Create citations and confidence signals to speed human review and enable safe handoffs.
  • Collaborate with product and domain experts to turn vague quality goals into measurable targets.
  • Instrument production traffic and generate golden datasets for regression tests.
  • Compare models and prompts across OpenAI, Anthropic, and open-weight options.
  • Bootstrap quality signals with synthetic data for new features.
  • Write templates and docs to help the team run evals.

Conocimientos

LLM/Agent systems
API integration
Document processing
Evaluation frameworks
Prompt engineering

Herramientas

OpenAI SDK
Anthropic SDK
Vercel AI SDK
LangGraph
Temporal Cloud

Descripción del empleo

About FirstIgnite

FirstIgnite makes software for university tech transfer offices. Those are the people who take research

coming out of a university lab and get it patented, licensed, or spun out into a company.

The role

We're hiring a Senior AI Agent Engineer. You'll build the agents in our product, and you'll build the

evals that tell us whether each change made them better or worse.

The work is document-heavy rather than chat. The agents run multi-step, call tools, read long and

inconsistently formatted source material, check it against existing records, and produce output that a

person reviews before anything happens with it.

Accuracy matters more here than speed or novelty. Most of the engineering effort goes into precision,

traceability, and getting the agent to hand off to a human at the right moment.

You'll report to the Head of Engineering and work with product and the full-stack team. If you've

shipped agents before, you've probably had the experience of changing a prompt and having no idea

whether you improved anything. That problem is most of this job.

What you'll do
  • Design and ship long-running, multi-step, tool-using agents on various AI SDKs and tooling, included but not limited to the OpenAI Agents SDK, the Anthropic Agent SDK, the Vercel AI SDK, LangGraph, MCP, and Temporal Cloud.
  • Wrap our APIs and our partners' APIs as tools an agent can call over MCP. Some of those systems are old, single-tenant, and outside our control, so a fair amount of the work is translation.
  • Get structured data out of long documents and match it against records that already exist. Expect entity resolution and fuzzy matching, and expect much of it to run in batch.
  • Stand up eval suites using various evaluation frameworks and tooling, included but not limited to Promptfoo, Braintrust, LangSmith, DeepEval, LLM-as-judge methods, and custom harnesses. Measure tool-use correctness, trajectory quality, and whether the agent finished the task.
  • Every agent here produces a draft that a person signs off on. Build the citations and confidence signals that make that review fast, and give the agent a clear way to escape.
  • Sit with product and domain experts and turn vague quality goals into something measurable. Sometimes the only dataset available for that is tiny, or confidential, or both.
  • Instrument production traffic, turn real customer interactions into golden datasets, and run them as regression tests.
  • Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt strategies and agent designs, and know what each option costs in latency and quality.
  • Bootstrap quality signal for features that have no production traffic yet. That usually means generating synthetic documents and test cases, including the ugly edge cases real customers will eventually send us, and knowing where synthetic data stops being a good proxy.
  • Write the templates, docs, and tooling the rest of the team needs to run evals without coming to you.
Required Qualifications
  • 3+ years of engineering experience, including hands-on work on LLM or agent systems that real users touched.
  • You've evaluated agents, not only models, and you know why single-turn accuracy says little about a multi-step run.
  • You've integrated against APIs you don't own, including old ones with bad documentation, and turned them into something an agent can call reliably.
  • You're comfortable with document pipelines: pulling data out, normalizing it, and checking it against a structured source of truth.
  • You've used at least one LLM evaluation framework, in-house tooling included.
  • You know how
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior AI Agent Engineer
Senior AI Agent Engineer

FirstIgnite • O Chazo

Presencial
EUR 90.000 - 130.000
Senior AI Agent Engineer
Senior AI Agent Engineer

FirstIgnite • Valencia

Presencial
EUR 60.000 - 90.000
Senior AI Agent Engineer
Senior AI Agent Engineer

FirstIgnite • Vitoria

Presencial
EUR 70.000 - 90.000
Senior Ai Agent Engineer
Senior Ai Agent Engineer

Firstignite • Arbo

Presencial
EUR 90.000 - 130.000
Senior AI Agent Engineer - Precision, Tooling & Evaluation
Senior AI Agent Engineer - Precision, Tooling & Evaluation

FirstIgnite • O Chazo

Presencial
EUR 90.000 - 130.000
Senior AI Agent Engineer - Build Precise, Multi-Step Agents
Senior AI Agent Engineer - Build Precise, Multi-Step Agents

FirstIgnite • Vitoria

Presencial
EUR 70.000 - 90.000
Senior AI Agent Engineer: Build & Evaluate Multi-Step Agents
Senior AI Agent Engineer: Build & Evaluate Multi-Step Agents

FirstIgnite • As Catro Rúas

Presencial
EUR 70.000 - 95.000
Senior AI Agent Engineer - Remote, Precision-Focused
Senior AI Agent Engineer - Remote, Precision-Focused

FirstIgnite • España

Presencial
EUR 90.000 - 130.000
Senior AI Agent Engineer: Build & Validate MultiStep Agents
Senior AI Agent Engineer: Build & Validate MultiStep Agents

FirstIgnite • Valencia

Presencial
EUR 60.000 - 90.000
AI Software Engineer | Spain
AI Software Engineer | Spain

Accenture España • Madrid

Presencial
EUR 90.000 - 130.000