Senior Site Reliability Engineer (Golang / Kubernetes)

Mirantis

Northern (KY)

Hybrid

USD 140,000 - 210,000

Full time

28 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Professional development
Conference attendance
Open-source collaboration

Job summary

Mirantis, an IREN company, seeks a Senior SRE to own SLIs/SLOs for the K0rdent Observability Framework (KOF) across hybrid, edge, and air-gapped deployments. You will derive meaningful SLI signals from Kubernetes, bare metal, and NVIDIA infrastructure and expose them through a clean API for Platform Administrators.

We expect you to define alerting and error-budget practices, partner with infrastructure, storage, and networking teams to instrument signals, and drive resolution of reliability

Qualifications

  • 5+ years in SRE, platform reliability, or related role.
  • Strong software engineering skills (Go or Python) with API/service experience.
  • Experience defining SLIs/SLOs and error budgets in production.
  • Hands-on with observability tools: metrics, logs, tracing.
  • Solid understanding of Kubernetes and related signals.
  • Excellent written and verbal communication across technical teams.

Responsibilities

  • Define SLIs and SLOs across the platform (Kubernetes, bare metal, NVIDIA).
  • Design API exposing SLIs and reliability state to admins and downstream systems.
  • Establish alerting and error-budget practices to maximize signal and minimize noise.
  • Partner with infra, storage, and networking teams to instrument and collect signals.
  • Diagnose reliability and performance issues across the observability stack and drive resolution.

Skills

SRE experience (5+ years)
Go or Python
APIs / services in production
Observability tooling
Kubernetes knowledge
Technical communication

Tools

Prometheus
VictoriaMetrics
OpenTelemetry
Grafana
K0rdent stack
Cluster API

Job description

  • Full-time
Company Description

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. https://www.mirantis.com/

Job Description

Define what reliability means for a GPU-accelerated AI platform and make it measurable. You will own the service-level indicators and objectives for the K0rdent Observability Framework (KOF) — deriving meaningful SLIs from the signals the platform already emits, and exposing them to Platform Administrators through a clean API. Work spans hybrid, edge, and air-gapped deployments built on the Mirantis K0rdent stack.

About the Role

We are looking for a Senior SRE who thinks past dashboards to the contract between a platform and its operators. The right candidate can look at raw telemetry from Kubernetes, bare metal, and NVIDIA infrastructure, decide which signals actually predict user-visible reliability, and turn them into SLIs and SLOs that operators can act on. You are equally comfortable writing the service that exposes those SLIs through an API and reasoning about error budgets, alerting quality, and signal-to-noise. You should be self-directed, able to own reliability definitions end to end, and effectively communicate them across teams.

Responsibilities

Define SLIs and SLOs based on the signals available across the platform — Kubernetes, bare-metal hosts, and NVIDIA infrastructure (BMC, InfiniBand, NVLink, UFM).

Design and build the API that exposes SLIs and reliability state to Platform Administrators and downstream systems.

Establish alerting and error-budget practices that maximize signal and minimize noise.

Partner with infrastructure, storage, and networking teams to ensure the right signals are instrumented and collected.

Diagnose reliability and performance issues across the observability stack and drive their resolution.

Qualifications

Required Qualifications

5+ years in SRE, platform reliability, or a closely related software/infrastructure role.

Strong software engineering skills (e.g., Go or Python) with experience building and operating APIs or services in production.

Demonstrated experience defining SLIs/SLOs and error budgets for real production systems.

Hands-on experience with observability tooling — metrics, logging, and tracing (e.g., Prometheus/VictoriaMetrics, OpenTelemetry, Grafana).

Solid understanding of Kubernetes and the signals it and its workloads emit.

Strong written and verbal communication with technical audiences.

Preferred

Experience instrumenting or monitoring bare-metal and NVIDIA infrastructure (BMC/Redfish, InfiniBand, NVLink, UFM).

Experience with the Mirantis K0rdent stack (K0rdent Enterprise, K0rdent AI, KOF) and Cluster API.

Familiarity with VictoriaMetrics/VictoriaLogs at scale.

Proven experience in sovereign or high-security air-gapped environments.

Additional Information

What does Mirantis offer you?

  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (Golang / Kubernetes)
Senior Site Reliability Engineer (Golang / Kubernetes)

JobCubby • Northern (KY)

Hybrid
USD 140,000 - 180,000
Competitive compensation package
Professional development
Conference attendance
+2
Senior Site Reliability Engineer (Golang / Kubernetes)
Senior Site Reliability Engineer (Golang / Kubernetes)

Mirantis • United States

Hybrid
USD 140,000 - 230,000
Competitive compensation package
Strong benefits plan
Professional development and training
+1
Senior Software Engineer (Golang) - remote in the US
Senior Software Engineer (Golang) - remote in the US

JobCubby • Northern (KY)

Hybrid
USD 140,000 - 210,000
Competitive compensation package
Strong benefits plan
Conferences and tech talks
AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

On-site
USD 120,000 - 180,000
Professional development and training
Conferences and working groups
Company outings and social events
+1
Senior SRE for GPU AI Platform | Go & Kubernetes
Senior SRE for GPU AI Platform | Go & Kubernetes

JobCubby • Northern (KY)

Hybrid
USD 140,000 - 180,000
Competitive compensation package
Professional development
Conference attendance
+2
Mirantis: Technical Product Marketer – K0rdent AI
Mirantis: Technical Product Marketer – K0rdent AI

Mosaec • Austin (TX)

On-site
USD 110,000 - 170,000
Competitive compensation package
Strong benefits plan
Open-source culture
Senior Kubernetes DevOps Engineer - K0rdent AI Apps/Core Services
Senior Kubernetes DevOps Engineer - K0rdent AI Apps/Core Services

Mirantis • San Jose (CA)

On-site
USD 120,000 - 160,000
Competitive compensation package
Professional development and training
Company outings and hackathons
Technical Product Marketer - K0rdent AI
Technical Product Marketer - K0rdent AI

Mirantis • United States

On-site
USD 120,000 - 180,000
Senior SRE: GPU AI Platform Reliability & SLIs
Senior SRE: GPU AI Platform Reliability & SLIs

Mirantis • Northern (KY)

Hybrid
USD 140,000 - 210,000
Competitive compensation
Professional development
Conference attendance
+1
Remote Technical Product Marketer, k0rdent AI – remote in the US
Remote Technical Product Marketer, k0rdent AI – remote in the US

Mirantis Inc. • Town of Florida (NY)

Hybrid
USD 120,000 - 180,000
Competitive compensation package
Strong benefits plan
Open-source innovation