Lead LLM Engineer

Licorne Society

Paris

Sur place

EUR 90 000 - 130 000

Plein temps

Il y a 3 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Passez les filtres ATS

Résumé du poste

Licorne Society is partnering with a fast-growing AI startup to hire a Lead LLM Engineer. You will own the architecture and quality of AI outputs, building end-to-end pipelines and evaluation frameworks that yield reliable, fast results in real workflows.

You will work on designing production systems, improving retrieval and tool usage, and iterating rapidly from real data. A strong focus on reliability and measurable impact guides every decision.

Qualifications

  • Proven experience shipping real LLM systems in production.
  • Strong understanding of RAG, tools, agents and structured outputs.
  • Ability to design full pipelines, not just prompts.
  • Experience building and using evaluation metrics and datasets.
  • Skilled at debugging retrieval, prompts, and model behavior.
  • Comfort with fast iteration from production data and logs.

Responsabilités

  • Design and evolve LLM/agent architecture.
  • Own output quality across key use cases (emails, document analysis, etc.).
  • Build evaluation systems with datasets and metrics.
  • Drive fast iteration loops from production data.
  • Improve retrieval, reasoning, and tool usage for reliability.
  • Ensure production reliability (latency, failure modes, fallback).
  • Collaborate with product and founders on what to build and why.

Connaissances

LLM systems
Evaluation-driven development
Debugging complex failures
Speed of iteration
System architecture
Reliability focus

Outils

Python (FastAPI)
Postgres
Google Cloud
LangGraph
LangChain
PostHog
Langfuse
Azure OpenAI

Description du poste

Licorne Society a été missionné par une startup IA en pleine croissance pour les aider à trouver leur Lead LLM Engineer.

What you will own

You will be responsible for one thing:
Make our AI outputs reliable, fast, and indispensable in real workflows.
Concretely:

  • Design and evolve our LLM / agent architecture
  • Own output quality across key use cases (emails, document analysis, etc.)
  • Build evaluation systems (datasets, metrics, regression detection)
  • Drive fast iteration loops from production data
  • Improve retrieval, reasoning, and tool usage
  • Ensure production reliability (latency, failure modes, fallback)
  • Work directly with product + founders on what to build and why
What this role is really about

Most teams fail because:

  • they don’t know what “good output” means
  • they don’t have evals
  • they iterate randomly
  • they overuse agents

Your job is to fix that.
You will turn:

  • vague user problems
  • into structured AI systems
  • with measurable performance
  • that improve every week
What you need to be excellent at
1. Shipping real LLM systems
  • You’ve built systems used in production (not demos)
  • You understand RAG, tools, agents, structured outputs
  • You can design full pipelines, not just prompts
2. Evaluation-driven development
  • You know how to define quality metrics
  • You build datasets from real usage
  • You run continuous evals to prevent regressions
3. Debugging complex failures
  • You can trace issues across:
    • retrieval
    • prompts
    • model behavior
  • You don’t guess — you isolate and fix
4. Speed of iteration
  • You move from problem → improvement in hours or days, not weeks
  • You use logs, traces, and data — not intuition alone
5. Strong judgment
  • You know when to:
    • use an agent vs a pipeline
    • add complexity vs simplify
  • You optimize for reliability and user value, not novelty
What we don’t care about
  • Number of years of experience
  • Whether you’ve used a specific framework
  • Fancy research credentials

If you can build, debug, and improve real systems, you’re a fit.

What success looks like (first 90 days)
  • Clear eval framework for core use cases
  • Measurable improvement in output quality
  • Faster iteration cycles across the team
  • Reduced hallucinations / failures
  • Stronger system architecture decisions
Stack (context, not requirements)
  • Python (FastAPI)
  • Postgres
  • Google Cloud
  • LangGraph / LangChain (evolving)
  • PostHog (product analytics)
  • Langfuse (LLM traces)
  • LLM APIs (Azure OpenAI)
Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

LLM Engineer
LLM Engineer

Licorne Society • Paris

Sur place
EUR 90 000 - 130 000
Senior Applied AI Engineer
Senior Applied AI Engineer

lemlist • Paris

Sur place
EUR 120 000 - 180 000
Consultant AI Software Engineer
Consultant AI Software Engineer

Jobtailor • Grenoble

Sur place
EUR 60 000 - 90 000
AI engineer
AI engineer

Lucis • Paris

Sur place
EUR 70 000 - 90 000
Relocation support
Open to frequent collaboration
AI engineering
AI engineering

Lucis (YC X25) • Paris

Sur place
EUR 70 000 - 90 000
Relocation support
Collaborative work environment
ML Infrastructure Engineer
ML Infrastructure Engineer

Whitecircle • Paris

Sur place
EUR 90 000 - 140 000
Engineering Manager | Applied AI
Engineering Manager | Applied AI

Jupus • Cologne

À distance
EUR 120 000 - 180 000
Fully Remote
MacBook + accessories
Urban Sports Club
Research Engineer, Forge
Research Engineer, Forge

Jobtailor • Paris

Sur place
EUR 70 000 - 100 000
Lead AI Engineer
Lead AI Engineer

Jobtailor • Paris

Sur place
EUR 90 000 - 140 000
Senior Applied AI Engineer
Senior Applied AI Engineer

lemlist • France

Sur place
EUR 90 000 - 130 000