Senior Site Reliability Engineer

Mirantis

Hyderabad

On-site

INR 3,500,000 - 6,000,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive compensation package
Professional development and training
Conferences and working groups
Hackathons and tech talks

Job summary

Mirantis is seeking a senior Kubernetes-focused DevOps/SRE engineer to own a customer-facing production region of a multi-tenant control plane for enterprise GPU infrastructure. You will build and run environments, pipelines, and tooling used by engineering teams across US, Europe, and APAC, ensuring fast, reproducible deployments from commit to release.

You will lead on-call ownership, harden Helm and CI/CD paths, and provide expert guidance on deployment topology, RBAC, and security.

Qualifications

  • 10+ years in DevOps/SRE with production ownership of Kubernetes environments.
  • Expert-level Kubernetes: workloads, networking, storage, RBAC, upgrades.
  • Proven incident response under SLA pressure and on-call rotations.
  • Strong CI/CD engineering: pipelines as code, GitHub Actions or equivalent.
  • Scripting and Go familiarity to read service code and trace failures.
  • Experience running stateful stacks in Kubernetes (PostgreSQL, identity providers, gateways).
  • Clear runbooks and incident write-ups across global time zones.

Responsibilities

  • Own the Kubernetes footprint across development, CI, pre-production, and one customer-facing region.
  • Operate production region to meet defined SLOs: capacity, upgrades, backups, DR drills, on-call.
  • Lead incident response: detection, mitigation, root-cause, blameless postmortems.
  • Build and maintain Helm charts and umbrella releases for platform services.
  • Own end-to-end CI/CD pipelines: build, test, image publish, release, hotfix flows.
  • Automate environment bootstrap and seeding for full-stack deployments.
  • Operate and troubleshoot supporting stack: PostgreSQL, Temporal, Keycloak, API gateway, brokers, observability.
  • Develop observability: metrics, dashboards, alerts, logs for dev and prod.
  • Consult with product teams on deployment topology, GPU scheduling, RBAC, networking, upgrades.
  • Enforce security and tenant isolation: least-privilege access, secrets, TLS, audits.
  • Drive infrastructure as code and repeatability; avoid snowflake environments.
  • Mentor engineers on Kubernetes and operational practices.

Skills

Kubernetes
SRE
CI/CD
Helm
Go
Terraform
Argo CD
Observability
On-call

Education

Bachelor's degree in Computer Science or related field

Tools

GitHub Actions
Docker
Helm
Terraform
Argo CD
Flux
Prometheus
Grafana

Job description

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment - on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. https://www.mirantis.com/

Job Description

We are looking for a senior Kubernetes-focused DevOps/SRE engineer to own both the developer platform and a customer-facing production region of a multi-tenant control plane for enterprise GPU infrastructure. You will build and run the environments, pipelines, and infrastructure tooling that our engineering teams across the US, Europe, and APAC depend on to ship daily.

This role spans both sides of the line. You will make our development, test, and pre-production clusters fast and reproducible, harden the Helm and CI/CD path from commit to release, and carry operational ownership - including on-call - for one of our smaller customer-facing production regions, under real availability commitments. That production experience makes you the internal expert on how k0rdent AI is deployed and operated - the person other teams consult, including the teams running our larger regions. Working within an agile framework, you will directly shape how quickly and safely changes reach production, and be accountable for how they behave once there.

Main Responsibilities:

Own the Kubernetes footprint across development, CI, pre-production, and one customer-facing production region - local kind clusters, shared dev and QA environments, and multi-cluster/multi-region topologies.

Operate your production region against defined SLOs: capacity and upgrade planning, patching, backup and restore, disaster recovery drills, and participation in an on-call rotation.

Lead incident response for your region - detection, mitigation, customer-impact assessment, root-cause analysis, and blameless postmortems that feed fixes back into the platform.

Build and maintain Helm charts and umbrella releases for the platform's services and dependencies, including versioning, values hygiene, and upgrade paths.

Own the CI/CD pipelines end to end - build, test, image publishing, chart packaging, release cutting, and hotfix/backport flows.

Automate environment bootstrap and seeding so any engineer can bring up a full stack - control plane, identity, gateway, database, workflow engine - with one command.

Operate and troubleshoot the supporting stack across test and production: PostgreSQL, Temporal, Keycloak, API gateway, message broker, and observability components.

Build observability and diagnostics - metrics, dashboards, alerting, log and audit access - that serve both engineering environments and production operations.

Consult with product teams and with the teams operating our larger regions on deployment topology, GPU and resource scheduling, RBAC, networking, and failure modes; validate upgrade and migration procedures and hand over runbooks.

Enforce security and tenant isolation in production: least-privilege access, secret handling, certificate and TLS lifecycle, image and dependency scanning, and audit evidence for compliance reviews.

Drive infrastructure as code and repeatability - no snowflake environments, no undocumented manual steps.

Mentor engineers on Kubernetes and operational practice, and raise the team's bar through review and documentation.

Qualifications

Required Skills/Abilities:

10+ years in DevOps, SRE, platform, or infrastructure engineering, including production ownership of customer-facing Kubernetes environments.

Expert-level Kubernetes: workloads, networking, storage, RBAC, resource management, CRDs and operators, and cluster upgrades - able to debug from kubectl and cluster internals rather than dashboards alone.

Proven incident response under SLA pressure - on-call rotations, escalation paths, postmortems, and follow-through on corrective action.

Strong CI/CD engineering - pipelines as code, reproducible builds, artifact and release management (GitHub Actions or equivalent).

Solid scripting and automation ability, and enough Go familiarity to read service code, trace a failure into it, and file a precise bug.

Experience running the stateful supporting stack - relational databases, identity providers, gateways, and message brokers - in Kubernetes, including backup, restore, and upgrade.

Track record as a technical consultant to other engineering teams: clear runbooks, design feedback, and incident write-ups across global time zones (strong written English).

Must Have

We expect deep production experience in several of these, and the engineering fundamentals to learn the rest rapidly.

Kubernetes Native: Kubernetes at scale, Cluster API, controllers/operators, Docker, and Helm chart authoring and lifecycle management.

Delivery: GitHub Actions or equivalent CI/CD, container registries, versioned release and backport workflows.

Infrastructure as Code: Terraform, Ansible, or equivalent, plus GitOps tooling (Argo CD, Flux).

Identity & API Management: Keycloak and API gateway operation - routing, plugins, TLS, rate limiting.

Data & Messaging: PostgreSQL operations and migrations, plus streaming/message-broker platforms (Kafka or equivalent).

Observability: Prometheus, Grafana, centralized logging, and alerting tied to SLOs.

Cloud: AWS - networking, IAM, load balancing, and managed Kubernetes

Nice to Have

k0s or k0rdent ecosystem experience.

GPU infrastructure on Kubernetes - device plugins, node feature discovery, scheduling and sharing of accelerators.

Temporal operations - namespaces, workers, schema upgrades.

Multi-region topologies, service mesh, or cross-cluster networking.

Python for test harnesses and automation; experience with pytest-based E2E suites.

Load and performance testing of API platforms.

Policy enforcement (OPA/Kyverno), secret management, and pen-test remediation.

OpenTelemetry, distributed tracing, or formal SLO/error-budget practice.

Compliance exposure - SOC 2, ISO 27001, or similar audit support.

CNCF open-source contributions.

Education and Experience:

Bachelor’s degree in Computer Science & Engineering or related field or 10 years related experience.

Additional Information

What does Mirantis offer you?

  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Software Engineer, Backend (Go)
Principal Software Engineer, Backend (Go)

Mirantis • Hyderabad

On-site
INR 4,000,000 - 8,000,000
Senior Data Platform Engineer
Senior Data Platform Engineer

SmartRecruiters, Inc. • India

Remote
INR 4,000,000 - 7,000,000
Competitive compensation package with强
Professional development and training
Attend conferences and working groups
+1
Senior Full-Stack TypeScript Engineer
Senior Full-Stack TypeScript Engineer

Mirantis • India

Remote
INR 3,000,000 - 5,000,000
Professional development
Conferences
Open-source culture
Technical Product Marketer - K0rdent AI
Technical Product Marketer - K0rdent AI

Embedded Shishya • Gopalganj

Hybrid
INR 11,483,000 - 15,311,000
Competitive compensation
Benefits plan
Senior SRE
Senior SRE

CloudRaft, Inc. • India

On-site
INR 2,500,000 - 5,000,000
Competitive salary
Premium health insurance
GPU infrastructure projects
Site Reliability Engineer(SRE)
Site Reliability Engineer(SRE)

CloudRaft, Inc. • India

On-site
INR 900,000 - 1,500,000
Competitive salary
Health insurance
GPU infrastructure exposure
+2
Senior Kubernetes Engineer (SME)
Senior Kubernetes Engineer (SME)

Operations • Mumbai

On-site
INR 1,200,000 - 2,400,000
Senior Devops Engineer
Senior Devops Engineer

Artech L.L.C. • India

On-site
INR 4,000,000 - 7,000,000
Senior Devops Engineer
Senior Devops Engineer

Randstad • Hyderabad

Hybrid
INR 4,200,000 - 6,500,000
Senior DevOps Engineer
Senior DevOps Engineer

Benchmarkit • Pune District

On-site
INR 1,400,000 - 2,400,000