Senior SRE: GPU AI Platform Reliability & SLIs

Mirantis

Northern (KY)

Hybrid

USD 140,000 - 210,000

Full time

35 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Professional development
Conference attendance
Open-source collaboration

Job summary

Mirantis, an IREN company, seeks a Senior SRE to own SLIs/SLOs for the K0rdent Observability Framework (KOF) across hybrid, edge, and air-gapped deployments. You will derive meaningful SLI signals from Kubernetes, bare metal, and NVIDIA infrastructure and expose them through a clean API for Platform Administrators.

We expect you to define alerting and error-budget practices, partner with infrastructure, storage, and networking teams to instrument signals, and drive resolution of reliability

Qualifications

  • 5+ years in SRE, platform reliability, or related role.
  • Strong software engineering skills (Go or Python) with API/service experience.
  • Experience defining SLIs/SLOs and error budgets in production.
  • Hands-on with observability tools: metrics, logs, tracing.
  • Solid understanding of Kubernetes and related signals.
  • Excellent written and verbal communication across technical teams.

Responsibilities

  • Define SLIs and SLOs across the platform (Kubernetes, bare metal, NVIDIA).
  • Design API exposing SLIs and reliability state to admins and downstream systems.
  • Establish alerting and error-budget practices to maximize signal and minimize noise.
  • Partner with infra, storage, and networking teams to instrument and collect signals.
  • Diagnose reliability and performance issues across the observability stack and drive resolution.

Skills

SRE experience (5+ years)
Go or Python
APIs / services in production
Observability tooling
Kubernetes knowledge
Technical communication

Tools

Prometheus
VictoriaMetrics
OpenTelemetry
Grafana
K0rdent stack
Cluster API

Job description

Mirantis, an IREN company, seeks a Senior SRE to own SLIs/SLOs for the K0rdent Observability Framework (KOF) across hybrid, edge, and air-gapped deployments. You will derive meaningful SLI signals from Kubernetes, bare metal, and NVIDIA infrastructure and expose them through a clean API for Platform Administrators.

We expect you to define alerting and error-budget practices, partner with infrastructure, storage, and networking teams to instrument signals, and drive resolution of reliability

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE for GPU AI Platform | Go & Kubernetes
Senior SRE for GPU AI Platform | Go & Kubernetes

JobCubby • Northern (KY)

Hybrid
USD 140,000 - 180,000
Competitive compensation package
Professional development
Conference attendance
+2
Senior SRE: GPU AI Platform Reliability & SLIs
Senior SRE: GPU AI Platform Reliability & SLIs

Mirantis • United States

Hybrid
USD 140,000 - 230,000
Competitive compensation package
Strong benefits plan
Professional development and training
+1
Senior Site Reliability Engineer (Golang / Kubernetes)
Senior Site Reliability Engineer (Golang / Kubernetes)

JobCubby • Northern (KY)

Hybrid
USD 140,000 - 180,000
Competitive compensation package
Professional development
Conference attendance
+2
Senior Site Reliability Engineer (Golang / Kubernetes)
Senior Site Reliability Engineer (Golang / Kubernetes)

Mirantis • United States

Hybrid
USD 140,000 - 230,000
Competitive compensation package
Strong benefits plan
Professional development and training
+1
Senior Site Reliability Engineer (Golang / Kubernetes)
Senior Site Reliability Engineer (Golang / Kubernetes)

Mirantis • Northern (KY)

Hybrid
USD 140,000 - 210,000
Competitive compensation
Professional development
Conference attendance
+1
Senior AI Cloud Deployment Engineer (SRE)
Senior AI Cloud Deployment Engineer (SRE)

Worky • Austin (TX)

On-site
USD 140,000 - 190,000
Competitive compensation
Strong benefits plan
Professional development
Senior SRE Lead - AI-Driven Reliability & Scale
Senior SRE Lead - AI-Driven Reliability & Scale

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model
Senior AIOps SRE for AI Data Center Platform
Senior AIOps SRE for AI Data Center Platform

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 148,000 - 276,000
Senior SRE: FinOps-Driven Infra, GPU & Scale
Senior SRE: FinOps-Driven Infra, GPU & Scale

Level AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1