Une candidature sur mesure pour ce poste — un CV et une lettre de motivation personnalisés qui correspondent à l’offre.
Jus Mundi is seeking a Senior AI engineer to own the applied AI systems behind our multilingual legal search and retrieval platform. You will lead autonomous agents, RAG pipelines, and evaluation harnesses to ensure accuracy and high quality answers grounded in sources.
You will design production-grade AI systems, fine-tune models with synthetic data, and mentor engineers while collaborating across squads to raise the bar in GenAI capabilities.
Jus Mundi is building the AI layer for international law and arbitration — turning a vast, multilingual corpus of case law, treaties, and legal analysis into answers practitioners can trust.
We are looking for a Senior AI engineer to own the applied AI systems behind that: search relevance and retrieval, RAG over legal documents, autonomous research agents, and the evaluation harnesses that keep them accurate. This is a hands‑on senior role with scope beyond a single squad.
You will take ambiguous problems — “make our search more relevant,” “let users ask a question and get a cited answer” — and turn them into concrete technical roadmaps, then ship them to production with measurable impact. The differentiators for this role are deep expertise in building autonomous agents (orchestration loops, tool calling, and memory — not just retrieval‑augmented systems), a critical command of evaluation and benchmarking, and hands‑on fine‑tuning and synthetic‑data generation. You will set patterns other engineers build on, and raise the bar through mentoring, reviews, and hiring.
AI depth is the core of the role; you should be comfortable integrating your work into our stack, but this is not a general full‑stack position.
Search & Retrieval Relevance: Own and improve retrieval and ranking over a large, multilingual legal corpus — a problem where relevance directly drives product value. Design, measure, and iterate on RAG pipelines that ground answers in the right sources.
Agent Orchestration (core of this role): Design and build agentic systems from the ground up — orchestration loops, tool/function calling, short- and long‑term memory, state management, planning, and multi‑step reasoning over documents. This goes well beyond wiring up a RAG pipeline: we need someone who has architected the control loop itself and made it reliable enough for a professional audience.
Evaluation & Benchmarking (core of this role): Design and run rigorous evals and benchmarks with a critical understanding of the methods themselves — knowing which metric measures what, where LLM‑as‑a‑judge is trustworthy and where it isn’t, and how to build regression suites and benchmark tracking that catch model degradation. Hold the bar on hallucinations and citation accuracy, which is non‑negotiable in a legal setting.
Fine‑Tuning & Synthetic Data: Fine‑tune open‑source or proprietary models (e.g. LoRA/QLoRA) when it beats prompting or retrieval, and generate high‑quality synthetic data to train and evaluate them. Curate domain datasets and make explicit trade-offs — precision vs. recall, cost vs. quality — rather than defaulting to the “perfect” academic solution.
Production & LLMOps: Deploy AI features to production and own their latency, cost, and reliability. Design systems that degrade gracefully under load or model failure instead of breaking catastrophically.
Leverage Across Squads: Build the shared tooling, libraries, and patterns (eval infra, retrieval components, agent scaffolding) that speed up other squads, not just your own — and pay down systemic AI/data debt before it stalls the team.
Technical Leadership: Act as the go‑to resource for complex modeling and GenAI problems. Level up engineers through code review and RFCs, and contribute to hiring — interviews, bar‑raising, and onboarding.