Site Reliability Engineer

Mistral.ai

Paris

Sur place

EUR 90 000 - 130 000

Plein temps

Il y a 5 jours
Soyez parmi les premiers à postuler
Générateur de candidature

N’envoyez pas un CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.

Passez les filtres ATS

Avantages offerts par ce poste

Healthcare coverage
Parental leave
Relocation support

Résumé du poste

Mistral.ai is seeking a Site Reliability Engineer (SRE) to join the Platform team in France. You will shape reliability, scalability, and performance of our AI platform and customer-facing apps, collaborating with software engineers and AI researchers to ensure high availability across web services, inference environments, and ML workloads.

This role balances daily production operations with long-term software engineering improvements, reducing toil and enabling seamless replication across HPC

Qualifications

  • Master’s degree in Computer Science, Engineering, or a related field.
  • 7+ years in a DevOps or SRE role with cloud and distributed systems expertise.
  • Hands-on experience with production reliability issues, in-production troubleshooting, and on-call rotations.
  • Strong proficiency with reliability KPIs, observability, alerting, and SLAs.
  • Experience with CI/CD, containerization, and orchestration tools such as Docker and Kubernetes.
  • Knowledge of monitoring/logging tooling (Prometheus, Grafana, ELK/Datadog).
  • Familiarity with IaC tools like Terraform or CloudFormation.
  • Scripting skills in Python, Go, or Bash; solid software development practices.

Responsabilités

  • Design, build, and maintain scalable, highly available infrastructure for web services and ML workloads.
  • Ensure production platforms are highly available; enable replication across HPC clusters.
  • Operate production systems and troubleshoot issues, including on-call responses and scaling.
  • Implement monitoring, alerting, and incident response to minimize downtime.
  • Develop and maintain CI/CD workflows, containerization, orchestration, logging, and monitoring tools.
  • Participate in on-call rotations and perform root cause analysis.
  • Drive automation to improve deployment, orchestration, and infrastructure reliability.
  • Collaborate with AI/ML researchers to enable reproducible experiments.
  • Build cloud-agnostic platforms to simplify infrastructure for science teams.
  • Work with security to meet best practices and compliance requirements.
  • Document processes to ensure knowledge sharing.

Connaissances

DevOps experience
SRE experience
Cloud computing
Distributed systems
CI/CD
Containerization
Kubernetes
Terraform
Python
Go/Bash scripting

Formation

Master’s degree in Computer Science/Engineering

Outils

Docker
Kubernetes
Terraform
CloudFormation
Prometheus
Grafana
ELK Stack
Datadog

Description du poste

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role

As a Site Reliability Engineer (SRE) on the Platform team, you will shape the reliability, scalability, and performance of our platform and customer-facing applications. You’ll work closely with software engineers and research teams to ensure our systems meet and exceed the expectations of both internal and external customers.

This role balances day-to-day operations on production systems with long-term software engineering improvements. Your work will reduce operational toil, foster reliability, and ensure high availability for our web services, inference environments, and ML workloads. You’ll enable seamless replication of work environments across multiple HPC clusters, directly impacting the stability and efficiency of our AI platform.

What You Will Do
  • Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support web services and ML workloads.

  • Ensure our platform, inference, and model training environments are always highly available and enable seamless replication across HPC clusters.

  • Operate systems and troubleshoot issues in production, including interrupts, on-call responses, and infrastructure scaling.

  • Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.

  • Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.

  • Participate in on-call rotations to respond to incidents and perform root cause analysis.

  • Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, and Terraform.

  • Collaborate with AI/ML researchers to enable safe and reproducible model-training experiments.

  • Build a cloud-agnostic platform that abstracts infrastructure complexities for science and engineering teams.

  • Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance.

  • Work with the security team to ensure infrastructure adheres to best practices and compliance requirements.

  • Document processes and procedures to ensure consistency and knowledge sharing across the team.

What We’re Looking For
  • A Master’s degree in Computer Science, Engineering, or a related field.

  • 7+ years of experience in a DevOps or SRE role, with strong expertise in cloud computing and distributed systems.

  • Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.

  • Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs.

  • Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.

  • Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.

  • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.

  • Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices.

  • Solid grasp of networking, security, and system administration concepts.

  • Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment.

  • Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus.

What We Offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Find Jobs in France on Arbeitnow

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Site Reliability Engineer, Mistral Cloud
Site Reliability Engineer, Mistral Cloud

Mistral • Paris

Sur place
EUR 90 000 - 130 000
Mistral Cloud - Site Reliability Engineer
Mistral Cloud - Site Reliability Engineer

Mistral • Paris

Sur place
EUR 90 000 - 120 000
Health insurance
Sport allowance
Meal vouchers
+2
Applied AI Engineer, Site Reliability Engineer - EMEA
Applied AI Engineer, Site Reliability Engineer - EMEA

United States Digital Space LLC • Paris

Sur place
EUR 90 000 - 130 000
Healthcare coverage
Relocation support
Meal and transportation allowances
Mistral Cloud - Software Engineer, Managed Kubernetes
Mistral Cloud - Software Engineer, Managed Kubernetes

Mistral.ai • Paris

Hybride
EUR 90 000 - 130 000
Relocation support
Meal and transportation allowances
Healthcare coverage
Research Engineer, Data Infrastructure
Research Engineer, Data Infrastructure

Mistral.ai • Paris

Hybride
EUR 90 000 - 140 000
Research Platform Engineer
Research Platform Engineer

Mistral.ai • Paris

Hybride
EUR 90 000 - 130 000
Healthcare coverage
Relocation support
Wellness programs
Senior AI Platform Reliability Engineer
Senior AI Platform Reliability Engineer

Mistral.ai • Paris

Sur place
EUR 90 000 - 130 000
Healthcare coverage
Parental leave
Relocation support
Applied AI Engineer, ML Infrastructure Engineer / Devops - EMEA
Applied AI Engineer, ML Infrastructure Engineer / Devops - EMEA

Mistral.ai • Paris

Sur place
EUR 85 000 - 120 000
Healthcare coverage
Retirement plans
Relocation support
+1
Research Platform Engineer
Research Platform Engineer

Mistral • Paris

Sur place
EUR 90 000 - 150 000
Health insurance
Transportation allowance
Sport allowance
+3
Infrastructure Solution Architect
Infrastructure Solution Architect

Mistral.ai • Paris

Sur place
EUR 90 000 - 130 000