Senior SRE: GPU AI Platform Reliability & SLIs

Mirantis

United States

Hybrid

USD 140,000 - 230,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive compensation package
Strong benefits plan
Professional development and training
Conferences and hackathons

Job summary

Mirantis is seeking a Senior SRE to define reliable SLIs/SLOs for a GPU-accelerated AI platform and own the end-to-end reliability contract between the platform and its operators. The role spans Kubernetes, bare-metal, and NVIDIA infrastructure, building an API to surface SLIs for Platform Administrators and downstream systems.

The candidate should craft alerting strategies, manage error budgets, and collaborate across infra, storage, and networking to ensure actionable signals and low noise.

Qualifications

  • 5+ years in SRE, platform reliability, or related role.
  • Strong software engineering skills (Go or Python) with production APIs/services experience.
  • Experience defining SLIs/SLOs and error budgets for real systems.
  • Hands-on observability tooling experience (metrics, logs, tracing).
  • Solid understanding of Kubernetes and its signals.

Responsibilities

  • Define SLIs and SLOs across the platform using available signals.
  • Design and build the API exposing SLIs and reliability state to admins and downstream systems.
  • Establish alerting and error-budget practices to maximize signal-to-noise.
  • Collaborate with infra, storage, and networking teams to instrument signals.
  • Diagnose reliability and performance issues across the observability stack.

Skills

Go
Python
APIs
Kubernetes
Observability
Prometheus
Grafana

Tools

Prometheus
VictoriaMetrics
OpenTelemetry
Grafana

Job description

Mirantis is seeking a Senior SRE to define reliable SLIs/SLOs for a GPU-accelerated AI platform and own the end-to-end reliability contract between the platform and its operators. The role spans Kubernetes, bare-metal, and NVIDIA infrastructure, building an API to surface SLIs for Platform Administrators and downstream systems.

The candidate should craft alerting strategies, manage error budgets, and collaborate across infra, storage, and networking to ensure actionable signals and low noise.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: GPU AI Platform Reliability & SLIs
Senior SRE: GPU AI Platform Reliability & SLIs

Mirantis • Northern (KY)

Hybrid
USD 140,000 - 210,000
Competitive compensation
Professional development
Conference attendance
+1
Senior SRE for GPU AI Platform | Go & Kubernetes
Senior SRE for GPU AI Platform | Go & Kubernetes

JobCubby • Northern (KY)

Hybrid
USD 140,000 - 180,000
Competitive compensation package
Professional development
Conference attendance
+2
Senior SRE: Lead Reliable AI Platform & Mentor the Team
Senior SRE: Lead Reliable AI Platform & Mentor the Team

Nscale • Houston (TX)

On-site
USD 170,000 - 265,000
Competitive base + equity
Real ownership from the start
Flexible work culture
Lead AI Platform Reliability Architect
Lead AI Platform Reliability Architect

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Platform SRE: Own Automation for AI/GPU Ops
Platform SRE: Own Automation for AI/GPU Ops

Nscale • Seattle (WA)

On-site
USD 130,000 - 200,000
Competitive base + equity
Real scope early
Flexible work style
Senior SRE Lead - AI-Driven Reliability & Scale
Senior SRE Lead - AI-Driven Reliability & Scale

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model
Senior Site Reliability Engineer (Golang / Kubernetes)
Senior Site Reliability Engineer (Golang / Kubernetes)

JobCubby • Northern (KY)

Hybrid
USD 140,000 - 180,000
Competitive compensation package
Professional development
Conference attendance
+2
Senior AI Cloud Deployment Engineer (SRE)
Senior AI Cloud Deployment Engineer (SRE)

Worky • Austin (TX)

On-site
USD 140,000 - 190,000
Competitive compensation
Strong benefits plan
Professional development
Senior SRE: AI/GPU Scale & Automation
Senior SRE: AI/GPU Scale & Automation

Nscale • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work expectations
Senior GPU Infra Reliability Engineer - Remote
Senior GPU Infra Reliability Engineer - Remote

Luma AI • United States

Remote
USD 180,000 - 240,000