Senior SRE: Scalable AI Platform & HPC

Mistral Ai

New York (NY)

On-site

USD 150,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Mistral AI is seeking an experienced Site Reliability Engineer to shape the reliability, scalability, and performance of our platform and customer-facing applications. You will work closely with software engineers and research teams to meet internal and external expectations.

The role focuses on designing scalable infrastructure, implementing robust monitoring, and driving CI/CD and automation to support high-availability workloads across HPC clusters and various environments.

Qualifications

  • Master’s degree in Computer Science, Engineering or related field is required.
  • 7+ years of experience in DevOps/SRE or equivalent roles.
  • Strong experience with cloud computing and highly available distributed systems.
  • Experience handling reliability KPIs (observability, monitoring, SLAs).
  • Hands-on with CI/CD, containerization and orchestration tools.

Responsibilities

  • Design, build, and maintain scalable, highly available infrastructure for web services and ML workloads.
  • Ensure production environments are resilient and on-call capable.
  • Improve monitoring, alerting, and incident response systems to minimize downtime.
  • Develop automation and workflows (CI/CD, containers, orchestration, dashboards).
  • Collaborate with AI/ML teams to enable reproducible experiments and safe model training.
  • Document processes for knowledge sharing and contribute to open-source projects.

Skills

CI/CD
Observability
Automation
Networking & security
SRE best practices
Scripting (Python/Go/Bash)
On-call experience
Problem solving
Startup environment
Communication

Education

Master’s degree in Computer Science or Engineering

Tools

Docker
Kubernetes
Terraform
CloudFormation
Prometheus
Grafana
ELK Stack
Datadog

Job description

Mistral AI is seeking an experienced Site Reliability Engineer to shape the reliability, scalability, and performance of our platform and customer-facing applications. You will work closely with software engineers and research teams to meet internal and external expectations.

The role focuses on designing scalable infrastructure, implementing robust monitoring, and driving CI/CD and automation to support high-availability workloads across HPC clusters and various environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Platform SRE for Scalable AI Systems
Senior Cloud Platform SRE for Scalable AI Systems

Mistral AI • Germany (OH)

On-site
USD 130,000 - 210,000
Healthcare coverage
Relocation support
Retirement plans
+2
Site Reliability Engineer, Mistral Cloud
Site Reliability Engineer, Mistral Cloud

Mistral AI • Germany (OH)

On-site
USD 130,000 - 210,000
Healthcare coverage
Relocation support
Retirement plans
+2
Senior SRE: AI-Driven Platform Reliability & Scale
Senior SRE: AI-Driven Platform Reliability & Scale

Medallia • McLean (VA)

On-site
USD 129,000 - 190,000
Health benefits
401(k) matching
Paid parental leave
+1
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior AI Cloud SRE — HPC & GPU Infrastructure
Senior AI Cloud SRE — HPC & GPU Infrastructure

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior SRE — Scale AI Systems, Equity Eligible
Senior SRE — Scale AI Systems, Equity Eligible

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior SRE: GPU AI Platform Reliability & SLIs
Senior SRE: GPU AI Platform Reliability & SLIs

Mirantis • United States

Hybrid
USD 140,000 - 230,000
Competitive compensation package
Strong benefits plan
Professional development and training
+1