Senior AI Engineer — GenAI & Autonomous Agents

Jus Mundi Partnership

Paris

Hybride

EUR 90 000 - 130 000

Plein temps

14 jours+
Générateur de candidature

Une candidature sur mesure pour ce poste — un CV et une lettre de motivation personnalisés qui correspondent à l’offre.

Passez les filtres ATS

Résumé du poste

Jus Mundi is seeking a Senior AI engineer to own the applied AI systems behind our multilingual legal search and retrieval platform. You will lead autonomous agents, RAG pipelines, and evaluation harnesses to ensure accuracy and high quality answers grounded in sources.

You will design production-grade AI systems, fine-tune models with synthetic data, and mentor engineers while collaborating across squads to raise the bar in GenAI capabilities.

Qualifications

  • Proven track record shipping LLM-powered features to production with measurable impact.
  • Strong command of vector databases, retrieval and ranking, and RAG architecture.
  • Hands-on expertise designing and building agent orchestration loops with memory, planning, and state management.

Responsabilités

  • Own retrieval and ranking improvements over a large multilingual legal corpus with measurable product impact.
  • Design and build agent orchestration systems including tool calling, memory, planning, and multi-step reasoning over documents.
  • Design and run rigorous evals and benchmarks to ensure accuracy and reduce hallucinations and mis-citations.
  • Fine-tune models (LoRA/QLoRA) and generate synthetic data to train/evaluate them, balancing precision vs recall and cost.
  • Deploy AI features to production, manage latency, cost, and reliability, with graceful degradation under load.
  • Build shared tooling and patterns to accelerate other squads and reduce AI/data debt.
  • Provide technical leadership, review code, guide RFCs, and contribute to hiring and onboarding.

Connaissances

Applied AI depth
Retrieval & RAG
Agent orchestration
Evaluation & benchmarking
Fine-tuning & synthetic data
Python for AI
System design
Mentoring & leadership
Multilingual NLP
Cloud infrastructure

Outils

OpenAI API
Anthropic
Mistral
Hugging Face
LangGraph/LangChain/LlamaIndex/AutoGen

Description du poste

About the Role

Jus Mundi is building the AI layer for international law and arbitration — turning a vast, multilingual corpus of case law, treaties, and legal analysis into answers practitioners can trust.

We are looking for a Senior AI engineer to own the applied AI systems behind that: search relevance and retrieval, RAG over legal documents, autonomous research agents, and the evaluation harnesses that keep them accurate. This is a hands‑on senior role with scope beyond a single squad.

You will take ambiguous problems — “make our search more relevant,” “let users ask a question and get a cited answer” — and turn them into concrete technical roadmaps, then ship them to production with measurable impact. The differentiators for this role are deep expertise in building autonomous agents (orchestration loops, tool calling, and memory — not just retrieval‑augmented systems), a critical command of evaluation and benchmarking, and hands‑on fine‑tuning and synthetic‑data generation. You will set patterns other engineers build on, and raise the bar through mentoring, reviews, and hiring.

AI depth is the core of the role; you should be comfortable integrating your work into our stack, but this is not a general full‑stack position.

Your mission at Jus Mundi
  • Search & Retrieval Relevance: Own and improve retrieval and ranking over a large, multilingual legal corpus — a problem where relevance directly drives product value. Design, measure, and iterate on RAG pipelines that ground answers in the right sources.

  • Agent Orchestration (core of this role): Design and build agentic systems from the ground up — orchestration loops, tool/function calling, short- and long‑term memory, state management, planning, and multi‑step reasoning over documents. This goes well beyond wiring up a RAG pipeline: we need someone who has architected the control loop itself and made it reliable enough for a professional audience.

  • Evaluation & Benchmarking (core of this role): Design and run rigorous evals and benchmarks with a critical understanding of the methods themselves — knowing which metric measures what, where LLM‑as‑a‑judge is trustworthy and where it isn’t, and how to build regression suites and benchmark tracking that catch model degradation. Hold the bar on hallucinations and citation accuracy, which is non‑negotiable in a legal setting.

  • Fine‑Tuning & Synthetic Data: Fine‑tune open‑source or proprietary models (e.g. LoRA/QLoRA) when it beats prompting or retrieval, and generate high‑quality synthetic data to train and evaluate them. Curate domain datasets and make explicit trade-offs — precision vs. recall, cost vs. quality — rather than defaulting to the “perfect” academic solution.

  • Production & LLMOps: Deploy AI features to production and own their latency, cost, and reliability. Design systems that degrade gracefully under load or model failure instead of breaking catastrophically.

  • Leverage Across Squads: Build the shared tooling, libraries, and patterns (eval infra, retrieval components, agent scaffolding) that speed up other squads, not just your own — and pay down systemic AI/data debt before it stalls the team.

  • Technical Leadership: Act as the go‑to resource for complex modeling and GenAI problems. Level up engineers through code review and RFCs, and contribute to hiring — interviews, bar‑raising, and onboarding.

Preferred Experience and Skills
  • Applied AI Depth: Proven track record shipping LLM-powered features to production — API integration (OpenAI, Anthropic, Mistral) and open‑source models (Hugging Face) — with clear, measurable impact.
  • Retrieval & RAG: Strong command of vector databases, retrieval and ranking, and RAG architecture, including how to actually improve relevance and evaluate it.
  • Agentic Systems (critical requirement): Deep, hands‑on expertise designing and building agent orchestration loops — tool/function calling, memory, planning, error recovery, and state management. You must be able to point to something you have actually built and shipped in this space (custom architectures or frameworks like LangGraph, LangChain, LlamaIndex, AutoGen). This is more important to us than RAG experience alone; strong RAG skills without proven agent‑building experience will not be sufficient.
  • Evaluation & Benchmarking (critical requirement): Deep, critical understanding of evaluation methodology — designing benchmarks, choosing and interpreting metrics, and building automated eval pipelines that reliably drive accuracy improvements. You should be able to reason about the limits of a given eval method, not just run one.
  • Fine‑Tuning & Synthetic Data (critical requirement): Demonstrated experience fine‑tuning models (LoRA/QLoRA), preparing training data, and generating synthetic data — plus the judgment to know when fine‑tuning is and isn’t worth it.
  • Engineering Foundations: Strong Python for AI and backend work, with enough full‑stack fluency (TypeScript/Node, some React) to integrate AI into product surfaces end to end.
  • Systems Judgment: Scalable API design and a habit of building for graceful degradation, testability, and reproducibility.
  • Senior Behaviours: Turns vague business problems into technical roadmaps in collaboration with product, design, and stakeholders; communicates results clearly to non‑technical audiences; mentors peers; and takes accountability for outcomes.
  • Nice to Have: Experience in legal‑tech or another high‑stakes, accuracy‑critical domain.
  • Multilingual NLP / retrieval experience.
  • Model inference optimization (vLLM, TensorRT‑LLM).
  • Cloud infrastructure (AWS, GCP, or Azure) and containerization.
  • Contributions to open‑source AI projects or a strong portfolio of shipped AI applications.
Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Senior Backend Engineer — Data Platform & AI Agents
Senior Backend Engineer — Data Platform & AI Agents

Business At Work • Paris

Sur place
EUR 90 000 - 130 000
Senior AI Engineer: GenAI & Autonomous Legal Agents
Senior AI Engineer: GenAI & Autonomous Legal Agents

Jus Mundi • Paris

Hybride
EUR 110 000 - 150 000
Hybrid working
Private health insurance
Restaurant vouchers
+5
Senior AI Engineer: GenAI & Autonomous Agents
Senior AI Engineer: GenAI & Autonomous Agents

Business At Work • Paris

Hybride
EUR 120 000 - 170 000
Hybrid working (2 days in office)
5 weeks vacation
Private health insurance
+2
Senior AI Engineer — GenAI & Autonomous Agents
Senior AI Engineer — GenAI & Autonomous Agents

Jus Mundi • Paris

Hybride
EUR 110 000 - 150 000
Hybrid working
Private health insurance
Restaurant vouchers
+5
Senior AI Engineer: Autonomous Agents & Legal GenAI
Senior AI Engineer: Autonomous Agents & Legal GenAI

Jus Mundi Partnership • Paris

Hybride
EUR 90 000 - 130 000
LLM Engineer
LLM Engineer

Leonar • Paris

Sur place
EUR 90 000 - 150 000
LLM Engineer
LLM Engineer

Licorne Society • Paris

Sur place
EUR 90 000 - 130 000
Lead LLM Engineer
Lead LLM Engineer

Licorne Society • Paris

Sur place
EUR 90 000 - 130 000
Senior AI Engineer (Agentic AI / AWS)
Senior AI Engineer (Agentic AI / AWS)

Gramian Consulting • France

Sur place
EUR 120 000 - 160 000
Remote work
EU work authorization
Lead AI Engineer
Lead AI Engineer

Jobtailor • Paris

Sur place
EUR 90 000 - 140 000