AI Engineer, Product

Mistral.ai

Paris

Sur place

EUR 70 000 - 100 000

Plein temps

Il y a 4 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Obtenez une réponse de cet employeur — un CV et une lettre de motivation adaptés exactement à ce qu’il recherche.

Passez les filtres ATS

Résumé du poste

Mistral.ai is seeking an AI quality engineer embedded in a product team to advance AI features across search, chat, documents, and audio. You will define what good looks like, measure it, run experiments, and ship improvements that affect quality, latency, safety, and reliability.

You will design evaluations, track key metrics, own prompts and system prompts, and run A/B tests to inform rollout decisions. Strong production ML experience and a product mindset are essential.

Qualifications

  • 3–4 years of experience in ML or software engineering with AI/ML production exposure.
  • Strong TypeScript or Python skills.
  • Production LLM experience: prompts, tool calls, system prompts.
  • Hands-on with evals and A/B testing; design metrics, not only run them.
  • Comfortable implementing directly in production code, not only notebooks.
  • Observability: logging, tracing, dashboards, alerts.
  • Form hypotheses, run experiments, interpret results, ship.
  • Clear communication, autonomous, production impact over experimentation for its own sake.

Responsabilités

  • Design and run evaluations for product area: reference tests, heuristics, model-graded checks for search relevance, chat quality, document understanding, or audio performance.
  • Define and track metrics: task success, helpfulness, hallucination proxies, safety signals, latency, cost.
  • Own prompt and orchestration design: write, test, and iterate prompts and system prompts.
  • Run A/B tests on prompts, models, configurations; analyze results and decide rollout/rollback.
  • Set up observability for LLM calls: logging, tracing, dashboards, alerts.
  • Manage model releases: canary and shadow traffic, sign-offs, SLO-based rollback criteria.
  • Improve core behaviors: memory policies, routing, tool-call reliability, or retrieval quality.
  • Create templates and documentation for evals and safe shipping.
  • Partner with Science to diagnose regressions and lead post-mortems.

Connaissances

TypeScript
Python
Evals & A/B testing
Observability
Product mindset

Description du poste

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role

Embedded directly in a product team as search, chat, documents, or audio, you'll improve AI-powered features through rigorous evaluation, prompt and orchestration design, and rapid experimentation. You'll own your domain's AI quality end-to-end: define what "good" looks like, measure it, run experiments, and ship what works. Work with Science to deliver measurable improvements to quality, latency, safety, and reliability.

What You Will Do
  • Design and run evaluations for your product area: reference tests, heuristics, model-graded checks tailored to search relevance, chat quality, document understanding, or audio performance.
  • Define and track metrics that matter: task success, helpfulness, hallucination proxies, safety flags, latency, cost.
  • Own prompt and orchestration design: write, test, and iterate on prompts and system prompts as a core part of your work.
  • Run A/B tests on prompts, models, and configurations; analyze results; make rollout or rollback decisions from data.
  • Set up observability for LLM calls: structured logging, tracing, dashboards, alerts.
  • Operate model releases: canary and shadow traffic, sign-offs, SLO-based rollback criteria, regression detection.
  • Improve core behaviors in your product area, whether that's memory policies, intent classification, routing, tool-call reliability, or retrieval quality.
  • Create templates and documentation so other teams can author evals and ship safely.
  • Partner with Science to diagnose regressions and lead post-mortems.
What We're Looking For
  • 3-4 years of experience; backgrounds that fit well include ML engineers moving closer to product, or software engineers with real AI/ML production experience.
  • Strong TypeScript or Python skills - we have both tracks depending on team fit.
  • Production LLM experience: prompts, tool/function calling, system prompts.
  • Hands-on with evals and A/B testing; you can design metrics, not just run them.
  • Comfortable implementing directly in product code, not only notebooks.
  • Observability experience: logging, tracing, dashboards, alerting.
  • Product mindset: form hypotheses, run experiments, interpret results, ship.
  • Clear communication, autonomous, and oriented toward production impact over experimentation for its own sake.
It would be ideal if you also have:
  • Safety systems experience: moderation, PII handling/redaction, guardrails.
  • Release operations: canary/shadowing, automated rollbacks, experiment platforms.
  • Prior work on search ranking, chat systems, document AI, or audio ML features.
What We Offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Find Jobs in France on Arbeitnow

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

AI Engineer, Product
AI Engineer, Product

Mistral • Paris

Sur place
EUR 70 000 - 110 000
Healthcare coverage
Relocation support
Retirement plans
+2
AI Engineer, Product
AI Engineer, Product

Mistral • Paris

Sur place
EUR 60 000 - 80 000
Competitive salary and equity package
Health insurance
Transportation allowance
+4
AI Engineer, Product
AI Engineer, Product

Mistral AI • Paris

Sur place
EUR 65 000 - 95 000
Applied AI Engineer, ML Infrastructure Engineer / Devops - EMEA
Applied AI Engineer, ML Infrastructure Engineer / Devops - EMEA

Mistral.ai • Paris

Sur place
EUR 85 000 - 120 000
Healthcare coverage
Retirement plans
Relocation support
+1
AI Scientist
AI Scientist

Mistral.ai • Paris

Sur place
EUR 90 000 - 130 000
Healthcare coverage
Relocation support
Wellness programs
+1
Engineering Team Lead, Backend
Engineering Team Lead, Backend

Mistral.ai • Paris

Sur place
EUR 90 000 - 130 000
Healthcare coverage
Relocation support
Meal & transportation allowances
Engineering Manager
Engineering Manager

Mistral.ai • Paris

Sur place
EUR 85 000 - 140 000
Applied AI, Forward Deployed Machine Learning Engineer - EMEA
Applied AI, Forward Deployed Machine Learning Engineer - EMEA

Mistral.ai • Paris

Sur place
EUR 70 000 - 100 000
Applied AI, Forward Deployed Machine Learning Engineer - EMEA
Applied AI, Forward Deployed Machine Learning Engineer - EMEA

Visa Hunt • Paris

Hybride
EUR 60 000 - 90 000
Infrastructure Solution Architect
Infrastructure Solution Architect

Mistral.ai • Paris

Sur place
EUR 90 000 - 130 000