Generative Ai Engineer: Agent-Driven Production

Wizeline

Salamanca

Presencial

EUR 60.000 - 90.000

Jornada completa

Hace 3 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo: un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Descripción de la vacante

Wizeline, a global AI-native technology solutions provider, is hiring a Senior AI Agent Engineer to design and ship long-running, tool-using agents and build evaluation suites. You will extract structured data from long documents, perform entity resolution, and ensure precise handoffs to humans when needed.

The role emphasizes accuracy, traceability, and collaboration with product and engineering teams. Remote work from Spain with a 3-month contract and potential extension, reporting to Head of

Formación

  • 3+ years of engineering experience with LLM or agent systems touched by real users.
  • Evaluated agents, not only models, understanding multi-step runs matter.
  • Integrated against external APIs and made them reliable for agents.
  • Comfortable with document pipelines: extract, normalize, verify vs source of truth.
  • Used at least one LLM evaluation framework (in-house tooling included).
  • Understanding LLM-as-judge biases (position/verbosity/judge drift) and strategies to mitigate.
  • Able to distinguish regression from noise and design experiments answering the right question.
  • Able to read customer transcripts, identify failures, and ship fixes with evals.
  • Writes clearly; engineers act on eval results they trust.
  • Based between NY ET and WEU time zones; Italy is the furthest east acceptable.

Responsabilidades

  • Design and ship long-running, multi-step, tool-using agents across AI SDKs and tooling.
  • Wrap internal and partner APIs as tools callable by agents.
  • Perform entity resolution and fuzzy matching on long documents against records.
  • Stand up eval suites and measure tool-use correctness and trajectory quality.
  • Generate citations and confidence signals to enable fast human reviews.
  • Collaborate with product and domain experts to turn vague goals into measurable outcomes.
  • Instrument production traffic and create golden datasets for regression tests.
  • Compare models and costs across OpenAI, Anthropic, and open-weight ecosystems.
  • Generate synthetic documents and test cases to bootstrap features with no traffic.
  • Create templates and tooling to enable the team to run evals independently.

Conocimientos

LLM experience
Agent systems
API integration
Document pipelines
LLM evaluation
Experiment design
Clear writing
Time-zone alignment

Educación

Bachelor's or Master's in CS/Engineering/Data Science

Herramientas

OpenAI Agents SDK
Anthropic Agent SDK
Vercel AI SDK
LangGraph
MCP
Temporal Cloud

Descripción del empleo

FirstIgnite makes software for university tech transfer offices. Those are the people who take research

coming out of a university lab and get it patented, licensed, or spun out into a company.

The role

We're hiring a Senior AI Agent Engineer. You'll build the agents in our product, and you'll build the

evals that tell us whether each change made them better or worse.

The work is document-heavy rather than chat. The agents run multi-step, call tools, read long and

inconsistently formatted source material, check it against existing records, and produce output that a

person reviews before anything happens with it.

Accuracy matters more here than speed or novelty. Most of the engineering effort goes into precision,

traceability, and getting the agent to hand off to a human at the right moment.

You’ll report to the Head of Engineering and work with product and the full-stack team. If you’ve

shipped agents before, you’ve probably had the experience of changing a prompt and having no idea

whether you improved anything. That problem is most of this job.

What you’ll do
  • Design and ship long-running, multi-step, tool-using agents on various AI SDKs and tooling, included but not limited to the OpenAI Agents SDK, the Anthropic Agent SDK, the Vercel AI SDK, LangGraph, MCP, and Temporal Cloud.
  • Wrap our APIs and our partners’ APIs as tools an agent can call over MCP. Some of those systems are old, single-tenant, and outside our control, so a fair amount of the work is translation.
  • Get structured data out of long documents and match it against records that already exist. Expect entity resolution and fuzzy matching, and expect much of it to run in batch.
  • Stand up eval suites using various evaluation frameworks and tooling, included but not limited to Promptfoo, Braintrust, LangSmith, DeepEval, LLM-as-judge methods, and custom harnesses. Measure tool-use correctness, trajectory quality, and whether the agent finished the task.
  • Every agent here produces a draft that a person signs off on. Build the citations and confidence signals that make that review fast, and give the agent a clear way to escape.
  • Sit with product and domain experts and turn vague quality goals into something measurable. Sometimes the only dataset available for that is tiny, or confidential, or both.
  • Instrument production traffic, turn real customer interactions into golden datasets, and run them as regression tests.
  • Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt strategies and agent designs, and know what each option costs in latency and quality.
  • Bootstrap quality signal for features that have no production traffic yet. That usually means generating synthetic documents and test cases, including the ugly edge cases real customers will eventually send us, and knowing where synthetic data stops being a good proxy.
  • Write the templates, docs, and tooling the rest of the team needs to run evals without coming to you.
Required Qualifications
  • 3+ years of engineering experience, including hands-on work on LLM or agent systems that real users touched.
  • You’ve evaluated agents, not only models, and you know why single-turn accuracy says little about a multi-step run.
  • You’ve integrated against APIs you don’t own, including old ones with bad documentation, and turned them into something an agent can call reliably.
  • You’re comfortable with document pipelines: pulling data out, normalizing it, and checking it against a structured source of truth.
  • You’ve used at least one LLM evaluation framework, in-house tooling included.
  • You know how LLM-as-judge methods break down (position bias, verbosity bias, judge drift) and what to do about it.
  • You can tell a real regression from noise, and design an experiment that answers the question being asked instead of a nearby one.
  • You can read a customer call transcript, work out which failures matter, and ship a fix and an eval for them.
  • You write clearly. Engineers won’t act on eval results they don’t read or don’t trust.
  • You’re based somewhere between New York time (ET) and Western European time. Italy is the furthest east we can go.
Preferred Qualifications
  • You’ve evaluated retrieval systems: RAG, hybrid search, reranking.
  • You’ve worked with agent orchestration frameworks like Temporal, LangGraph, or the OpenAI Agents SDK, and you know how long-running tool use goes wrong.
  • You have a background in information retrieval or search relevance.
  • You’ve worked somewhere an agent's output carried financial or compliance consequences.
  • You’ve built internal tooling that non-engineers used on their own to label and review model output.

This is a fully remote, full-time permanent position available to candidates located within the New York (ET) through Western Europe time zones, with flexible working hours to support collaboration across regions.

FirstIgnite makes software for university tech transfer offices. Those are the people who take research

coming out of a university lab and get it patented, licensed, or spun out into a company.

The role

We're hiring a Senior AI Agent Engineer. You'll build the agents in our product, and you'll build the

evals that tell us whether each change made them better or worse.

The work is document-heavy rather than chat. The agents run multi-step, call tools, read long and

inconsistently formatted source material, check it against existing records, and produce output that a

person reviews before anything happens with it.

Accuracy matters more here than speed or novelty. Most of the engineering effort goes into precision,

traceability, and getting the agent to hand off to a human at the right moment.

You’ll report to the Head of Engineering and work with product and the full-stack team. If you’ve

shipped agents before, you’ve probably had the experience of changing a prompt and having no idea

whether you improved anything. That problem is most of this job.

What you’ll do
  • Design and ship long-running, multi-step, tool-using agents on various AI SDKs and tooling, included but not limited to the OpenAI Agents SDK, the Anthropic Agent SDK, the Vercel AI SDK, LangGraph, MCP, and Temporal Cloud.
  • Wrap our APIs and our partners’ APIs as tools an agent can call over MCP. Some of those systems are old, single-tenant, and outside our control, so a fair amount of the work is translation.
  • Get structured data out of long documents and match it against records that already exist. Expect entity resolution and fuzzy matching, and expect much of it to run in batch.
  • Stand up eval suites using various evaluation frameworks and tooling, included but not limited to Promptfoo, Braintrust, LangSmith, DeepEval, LLM-as-judge methods, and custom harnesses. Measure tool-use correctness, trajectory quality, and whether the agent finished the task.
  • Every agent here produces a draft that a person signs off on. Build the citations and confidence signals that make that review fast, and give the agent a clear way to escape.
  • Sit with product and domain experts and turn vague quality goals into something measurable. Sometimes the only dataset available for that is tiny, or confidential, or both.
  • Instrument production traffic, turn real customer interactions into golden datasets, and run them as regression tests.
  • Compare models against each other (OpenAI, Anthropic, open-weight), along with prompt strategies and agent designs, and know what each option costs in latency and quality.
  • Bootstrap quality signal for features that have no production traffic yet. That usually means generating synthetic documents and test cases, including the ugly edge cases real customers will eventually send us, and knowing where synthetic data stops being a good proxy.
  • Write the templates, docs, and tooling the rest of the team needs to run evals without coming to you.
Required Qualifications
  • 3+ years of engineering experience, including hands-on work on LLM or agent systems that real users touched.
  • You’ve evaluated agents, not only models, and you know why single-turn accuracy says little about a multi-step run.
  • You’ve integrated against APIs you don’t own, including old ones with bad documentation, and turned them into something an agent can call reliably.
  • You’re comfortable with document pipelines: pulling data out, normalizing it, and checking it against a structured source of truth.
  • You’ve used at least one LLM evaluation framework, in-house tooling included.
  • You know how LLM-as-judge methods break down (position bias, verbosity bias, judge drift) and what to do about it.
  • You can tell a real regression from noise, and design an experiment that answers the question being asked instead of a nearby one.
  • You can read a customer call transcript, work out which failures matter, and ship a fix and an eval for them.
  • You write clearly. Engineers won’t act on eval results they don’t read or don’t trust.
  • You’re based somewhere between New York time (ET) and Western European time. Italy is the furthest east we can go.
Preferred Qualifications
  • You’ve evaluated retrieval systems: RAG, hybrid search, reranking.
  • You’ve worked with agent orchestration frameworks like Temporal, LangGraph, or the OpenAI Agents SDK, and you know how long-running tool use goes wrong.
  • You have a background in information retrieval or search relevance.
  • You’ve worked somewhere an agent's output carried financial or compliance consequences.
  • You’ve built internal tooling that non-engineers used on their own to label and review model output.

This is a fully remote, full-time permanent position available to candidates located within the New York (ET) through Western Europe time zones, with flexible working hours to support collaboration across regions.

  • GKE
  • AlloyDB

Wizeline, a global AI-native technology solutions provider, develops cutting-edge, AI-powered digital products and platforms. We partner with clients to leverage data and AI, accelerating market entry and driving business transformation. As a global community of innovators, we foster a culture of growth, collaboration, and impact.

With the right people and the right ideas, there’s no limit to what we can achieve

Sounds awesome, right? Now, let’s make sure you’re a good fit for the role:

Key Responsibilities
  • Collaborate with data scientists and engineers to orchestrate LLMs and tools into complex AI workflows.
  • Implement optimized vector storage and indexing systems for NLP.
  • Develop tools and frameworks for prompt management, automated evaluations, and observability.
  • Monitor and improve agentic systems performance for accuracy and efficiency.
  • Provide ongoing support, troubleshoot issues, and implement updates for ML solutions
  • Build and maintain data processing pipelines for high volumes of structured and unstructured data.
  • Integrate third-party AI APIs (internal and external) to extend the functionality of systems.
  • Stay updated on GenAI, NLP, ML, and IR technologies, incorporating best practices and leveraging cloud infrastructure for efficiency.
  • Build and extend internal products on top of the LangChain ecosystem (LangGraph, LangSmith) to support prompt management, evaluation, and agent orchestration across the team.
Must-have Skills
  • Bachelor's or Master's degree in Computer Science, Engineering, Data Science, or a related STEM field or equivalent work experience.
  • 4+ years of industrial experience in a machine learning engineering or data engineering role.
  • Strong programming skills in Python and/or another high-level language commonly used in machine learning.
  • Experience deploying LLMs, implementing automated evaluation pipelines (LLM-as-a-judge), and architecting multi-agent systems that utilize tool-calling and long-term memory to solve non-linear problems.
Nice-to-have
  • AI Tooling Proficiency : Leverage one or more AI tools to optimize and augment day-to-day work, including drafting, analysis, research, or process automation. Provide recommendations on effective AI use and identify opportunities to streamline workflows.
  • Familiarity with cloud-based infrastructure and services (e.g., AWS and GCP), Docker, and the Git version control system
  • Familiarity with consuming and integrating APIs in a reliable and secure manner.
What we offer
  • Commitment to Professional Development
  • Flexible and Collaborative Culture
  • Total Rewards

*Specific benefits are determined by the employment type and location.

Find out more about our culture here.

Details
  • Contract Duration: 3 months with extension
  • Remote work from Spain
Responsibilities
  • Design, develop, and support AI/ML solutions and distributed applications
  • Work with LLM architectures and AI-driven services in cloud-based environments
  • Develop and maintain Python-based applications, APIs, and microservices
  • Support integration, deployment, and optimization of AI/ML workflows
  • Collaborate with engineering and DevOps teams to ensure scalable and secure application delivery
  • Participate in CI/CD processes and DevOps practices for automated deployments
Skills
  • Minimum 2 years of hands-on experience with Python development
  • Good understanding of AI/ML concepts, LLM architectures, and distributed computing
  • Experience with microservices development and management
  • Knowledge of CI/CD and DevOps tools such as Git and Jenkins
  • Understanding of application security best practices
  • Familiarity with public cloud platforms, preferably AWS
  • Knowledge of vector and columnar databases
  • Understanding of Agile methodologies and collaborative development practices
  • AWS and/or DevOps certifications are considered a strong advantage

SNI sp. z o.o. will process personal data for the purpose of the recruitment process in accordance with Data Privacy Policy. The data may also be stored and processed for future recruitment purposes, in accordance with the given consent.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

AI Engineer
AI Engineer

Wizeline • España

Presencial
EUR 70.000 - 100.000
Professional Development
Flexible culture
Total Rewards
Senior Ai Agent Engineer
Senior Ai Agent Engineer

Firstignite • Arbo

Presencial
EUR 90.000 - 130.000
Senior AI Software Engineer
Senior AI Software Engineer

Wizeline • Barcelona

Híbrido
EUR 60.000 - 90.000
Health benefits
Retirement plans
Global mobility opportunities
+4
Senior Software Engineer – Fullstack
Senior Software Engineer – Fullstack

Luzia • Madrid

Presencial
EUR 45.000 - 70.000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Migx • Barcelona

Híbrido
EUR 90.000 - 130.000
Hybrid work model
25 holiday days per year
Career development opportunities
+1
Mid AI Engineer
Mid AI Engineer

Shalion • Barcelona

Híbrido
EUR 45.000 - 65.000
Flexible benefits
Office perks like fresh fruit and coffee
Dynamic and innovative work environment
Senior AI/ML Engineer
Senior AI/ML Engineer

SNI • España

A distancia
EUR 70.000 - 100.000
Lead AI Engineer (Agentic AI)
Lead AI Engineer (Agentic AI)

Visium SA • Barcelona

Híbrido
EUR 60.000 - 90.000
Competitive compensation package
Flexible working culture
Employee Stock Ownership Plan
AI Engineer
AI Engineer

Clarity • Madrid

Híbrido
EUR 70.000 - 105.000
Competitive pay
Location flexibility
Generous time off
+4
Lead AI Engineer
Lead AI Engineer

InteractiveAI Limited • Madrid

Híbrido
EUR 110.000 - 130.000
Equity plan
Health & wellness allowances
Private health insurance
+3