AI Cloud Platform SRE — Scale, Reliability & Uptime

Mistral

Paris

Sur place

EUR 90 000 - 130 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Résumé du poste

Mistral is seeking a Site Reliability Engineer to shape the reliability, scalability, and performance of our Cloud Platform and customer-facing applications. You will work closely with software engineers and product teams to ensure systems meet and exceed expectations for internal and external users.

This role is critical for maintaining stability of our infrastructure, enabling seamless experiences for users and developers.

Qualifications

  • Master's degree in Computer Science, Engineering, or a related field.
  • 5+ years in a DevOps or SRE role, with strong experience in bare metal infra and distributed systems.
  • Hands-on with in-production troubleshooting, root-cause analysis, and on-call rotations.
  • Experience with reliability KPIs, observability, alerting, and SLAs.
  • Proficiency with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.
  • Knowledge of monitoring/logging/alerting tools such as Prometheus, Grafana, ELK/Datadog.
  • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
  • Scripting skills (Python, Go, Bash) and solid software development fundamentals.
  • Strong networking, security, and system administration concepts.
  • Nice-to-have: AI/ML environment, HPC, or modern AI-oriented solutions.

Responsabilités

  • Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support our Cloud platform.
  • Operate systems in production, handle on-call responses, and scale infrastructure.
  • Implement and improve monitoring, alerting, and incident response to minimize downtime.
  • Develop and maintain CI/CD workflows, containerization, orchestration, logging, and monitoring tooling.
  • Participate in on-call rotations and perform RCAs to prevent recurrence.
  • Drive infrastructure automation and platform deployment improvements.
  • Collaborate with software engineers to enable safe, reproducible model-training experiments.
  • Develop cloud platforms that abstract infrastructure complexities for science and engineering teams.
  • Create workflows and tooling to improve reliability, availability, and performance.
  • Ensure security and compliance in collaboration with the security team.
  • Document processes to promote knowledge sharing.

Connaissances

DevOps
SRE
CI/CD
Docker
Kubernetes
Observability
Prometheus
Grafana
ELK
Terraform
Python
Networking

Formation

Master's degree in Computer Science or Engineering

Outils

Docker
Kubernetes
Terraform
CloudFormation
Datadog
ELK Stack
Datadog
Python

Description du poste

Mistral is seeking a Site Reliability Engineer to shape the reliability, scalability, and performance of our Cloud Platform and customer-facing applications. You will work closely with software engineers and product teams to ensure systems meet and exceed expectations for internal and external users.

This role is critical for maintaining stability of our infrastructure, enabling seamless experiences for users and developers.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

Site Reliability Engineer, Mistral Cloud
Site Reliability Engineer, Mistral Cloud

Mistral • Paris

Sur place
EUR 90 000 - 130 000
Backend Engineer - Cloud AI Platforms
Backend Engineer - Cloud AI Platforms

Mistral • Paris

Hybride
EUR 70 000 - 110 000
Data Infrastructure Engineer - Scale AI Training Platform
Data Infrastructure Engineer - Scale AI Training Platform

Mistral • Paris

Hybride
EUR 90 000 - 120 000
Healthcare coverage
Relocation support
Wellness programs
Senior Full-Stack Engineer, AI Cloud Platform
Senior Full-Stack Engineer, AI Cloud Platform

Mistral • Paris

Hybride
EUR 90 000 - 130 000
Platform SRE: Kubernetes, Azure, & CI/CD Reliability
Platform SRE: Kubernetes, Azure, & CI/CD Reliability

Allianz Partners • Saint-Ouen-sur-Seine

Sur place
EUR 70 000 - 110 000
Cloud Platform Engineering Lead – Build Scalable AI Systems
Cloud Platform Engineering Lead – Build Scalable AI Systems

Mistral • Paris

Sur place
EUR 90 000 - 130 000
Backend Engineer for AI Platform & APIs
Backend Engineer for AI Platform & APIs

Mistral • Paris

Sur place
EUR 70 000 - 110 000
Healthcare coverage
Parental leave
Retirement plans
+4
SRE - Reliability & Automation for Cloud Infra
SRE - Reliability & Automation for Cloud Infra

Scaleway • Toulouse

Hybride
EUR 70 000 - 110 000
Senior SRE - Scale Real-Time Data Infra
Senior SRE - Scale Real-Time Data Infra

Pigment • Paris

Sur place
EUR 70 000 - 100 000
Platform Security Engineer: Secure AI Infra & Pipelines
Platform Security Engineer: Secure AI Infra & Pipelines

Mistral • Paris

Hybride
EUR 90 000 - 130 000