Senior Cloud Platform SRE for Scalable AI Systems

Mistral AI

Germany (OH)

On-site

USD 130,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Healthcare coverage
Relocation support
Retirement plans
Wellness programs
Meal and transportation allowances

Job summary

Mistral AI is seeking a Site Reliability Engineer to join the Cloud Platform team. You will shape reliability, scalability, and performance for our cloud services and customer-facing apps, collaborating with software engineers and product teams to meet high standards of uptime and efficiency.

You will design and operate scalable infrastructure, improve monitoring, and drive automation across CI/CD, containers, and orchestration while ensuring security and compliance with internal policies.

Qualifications

  • Master's degree in Computer Science, Engineering, or a related field.
  • 5+ years of experience in DevOps or SRE roles with distributed systems context.
  • Hands-on experience with production incidents, on-call rotations, and RCA.
  • Proficiency in observability, alerting, and SLAs.
  • Experience with CI/CD, containerization, and orchestration.

Responsibilities

  • Design, build, and maintain scalable, highly available infrastructure for the Cloud Platform.
  • Operate systems in production, troubleshoot, and handle on-call rotations.
  • Improve monitoring, alerting, and incident response to minimize downtime.
  • Develop workflows/tools for CI/CD, containerization, orchestration, logging.
  • Collaborate with engineers to enable reproducible model-training experiments.
  • Build cloud platform abstractions to reduce infrastructure complexity for teams.
  • Document processes to ensure knowledge sharing across the team.

Skills

DevOps
SRE
Observability
On-call
Python
Go
Bash
Networking
Security
Team collaboration
AI/ML environments

Education

Master's degree in CS/Engineering

Tools

Docker
Kubernetes
CI/CD
Terraform
CloudFormation
Prometheus
Grafana
ELK Stack
Datadog

Job description

Mistral AI is seeking a Site Reliability Engineer to join the Cloud Platform team. You will shape reliability, scalability, and performance for our cloud services and customer-facing apps, collaborating with software engineers and product teams to meet high standards of uptime and efficiency.

You will design and operate scalable infrastructure, improve monitoring, and drive automation across CI/CD, containers, and orchestration while ensuring security and compliance with internal policies.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, Mistral Cloud
Site Reliability Engineer, Mistral Cloud

Mistral AI • Germany (OH)

On-site
USD 130,000 - 210,000
Healthcare coverage
Relocation support
Retirement plans
+2
Senior Site Reliability Engineer — NYC, Equity & Visa Support
Senior Site Reliability Engineer — NYC, Equity & Visa Support

Mistral • New York (NY)

Hybrid
USD 140,000 - 190,000
Competitive salary and equity
Healthcare: Medical/Dental/Vision for你
401K with match
+6
Senior SRE: AI-Driven Platform Reliability & Scale
Senior SRE: AI-Driven Platform Reliability & Scale

Medallia • McLean (VA)

On-site
USD 129,000 - 190,000
Health benefits
401(k) matching
Paid parental leave
+1
SRE, Cloud Platform – AI-Driven Reliability & Observability
SRE, Cloud Platform – AI-Driven Reliability & Observability

Visa • Austin (TX)

On-site
USD 88,000 - 137,000
Senior SRE: AI Cloud Reliability & Scale
Senior SRE: AI Cloud Reliability & Scale

Neura Market • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Health, dental, vision coverage
401k with company match
Wellness stipend
+1
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior SRE: AI Cloud Platform & Kubernetes Expert
Senior SRE: AI Cloud Platform & Kubernetes Expert

Lambda • Bellevue (WA)

On-site
USD 180,000 - 260,000
Health insurance
Dental insurance
Vision insurance
+3
Senior SRE: AI Cloud Platform & Kubernetes
Senior SRE: AI Cloud Platform & Kubernetes

Lambda Inc. • San Francisco (CA)

Hybrid
USD 190,000 - 270,000
Health insurance
401k with company match
Flexible PTO
+2
Senior Site Reliability Engineer - AI Cloud Platform
Senior Site Reliability Engineer - AI Cloud Platform

Lambda • United States

Hybrid
USD 160,000 - 220,000