Senior Site Reliability Engineer

Infinx

Bengaluru

On-site

INR 2,500,000 - 5,000,000

Full time

17 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Infinx, a healthcare payment solutions provider, Bengaluru-based, seeks a Senior Site Reliability Engineer to own the availability and performance of production systems across AWS and IBM Cloud. This is hands-on engineering, not a rotation, with automation, service levels, and architecture changes to prevent incidents.

You will work with Kubernetes on AWS EKS and OpenShift on IBM Cloud, define observability, IaC, and incident processes, ensuring one system with consistent SLIs and reliable

Qualifications

  • 4–6 years in site reliability engineering, DevOps, or production cloud infrastructure.
  • Hands-on with AWS and IBM Cloud across multi-cloud environments.
  • Strong knowledge of Kubernetes, Terraform, and OpenShift.
  • Experience with observability stacks and incident management.

Responsibilities

  • Own SLIs, SLOs, and error budgets for defined services with product owners.
  • Act as incident commander for high-severity events and drive blameless postmortems.
  • Plan capacity and headroom across cloud and on‑prem resources.
  • Define infrastructure as code and automate remediation across providers.

Skills

AWS
IBM Cloud
Kubernetes
Terraform
OpenShift
Python
Bash
Prometheus
Grafana
OpenTelemetry
Incident Management
Networking
Security & Compliance

Tools

Terraform
Kubernetes
OpenShift
Prometheus
Grafana
OpenTelemetry

Job description

Company Overview

Infinx Healthcare is a leading healthcare payment solutions provider specializing in next-generation, cloud-based SaaS products. Our mission is to maximize and preserve revenue across the US healthcare revenue cycle by combining human expertise with artificial intelligence.


Our platform carries critical clinical transactions for providers across the United States. Reliability engineering sits close to the centre of our engineering organisation rather than at its edge.


To learn more, visit www.infinx.com.



About the Role

We are looking for a Senior Site Reliability Engineer to own the availability and performance of production systems that span AWS, IBM Cloud, etc.. This is a hands‑on engineering role rather than a monitoring rotation. You will write the automation, define the service levels, and change the architecture that keeps recurring incidents from recurring.


Workloads run on managed Kubernetes in AWS and on Red Hat OpenShift in IBM Cloud. Your job is to make that estate behave like one system: consistent service levels, one observability pane, one incident process, and infrastructure defined as code no matter which provider sits underneath.



What You Will Own


Reliability and Service Levels


  • Own SLIs, SLOs, and error budgets for a defined set of critical services, negotiated with product and service owners rather than imposed on them.

  • Use error‑budget burn to make real prioritisation calls, including slowing a rollout when the budget is spent.

  • Run capacity and headroom planning across cloud and on‑premises resources, where procurement lead times differ by months.

  • Push reliability improvements upstream into application design reviews instead of absorbing them at the



Incident Response and Learning


  • Act as incident commander for high‑severity events, coordinating triage, escalation, customer communication, and resolution.

  • Facilitate blameless postmortems and make sure follow‑up actions are assigned, scheduled, and actually closed.

  • Reduce time to detection by improving signal quality. Every alert should be actionable, urgent, and owned by someone.

  • Publish reliability metrics and incident trends to engineering leadership, with honest commentary rather than



Hybrid Cloud and Platform Operations


  • Operate production Kubernetes across Amazon EKS and Red Hat OpenShift on IBM Cloud, including version upgrades, node lifecycle, and cluster hardening.

  • Own hybrid connectivity: AWS Direct Connect, IBM Cloud Direct Link, site‑to‑site VPN, transit routing, splithorizon DNS, and cross‑environment identity.

  • Maintain Terraform as the single source of truth across providers, covering module design, state isolation, drift detection, and safe promotion paths.

  • Design and prove failover and disaster‑recovery paths that cross provider boundaries, and validate them through

  • Advise on workload placement against cost, latency, data residency, and compliance constraints.



Observability and Automation


  • Build a federated observability layer that gives one view across both clouds, based on OpenTelemetry using

  • Standardise instrumentation so services emit comparable telemetry regardless of where they run.

  • Write production‑grade Python and Bash to remove toil through auto‑remediation, self‑service tooling, and safe

  • Set and defend a toil budget. If the team spends more than its target share of time on manual operations, reducing that becomes the work.



Security and Compliance


  • Operate within HIPAA and SOC 2 expectations: least‑privilege access, key management, audit trails, and evidence that holds up under review.

  • Contribute to cloud security posture management and vulnerability remediation across both cloud providers.

  • Enforce secure‑by‑default platform patterns including network segmentation, secrets management, admission control, and image provenance.



Technical Leadership


  • Mentor engineers on reliability practice and review their designs and change plans.

  • Write and maintain the standards others build against, and document decisions so they outlive any one

  • Represent reliability in architecture forums, vendor conversations, and audit discussions.



Must‑Have Skills

Experience


  • 4–6 years in site reliability engineering, DevOps, or production cloud infrastructure, including at least three years carrying on‑call responsibility for systems you helped build.



Hybrid and Multi‑Cloud


  • Deep, hands‑on production experience with AWS and preferably IBM Cloud.

  • Practical experience running workloads across more than one environment, including the networking and identity glue between them.



Containers and Platform


  • Production ownership of Kubernetes beyond day‑to‑day kubectl: upgrades, capacity, networking, storage, and

  • Docker and Helm. OpenShift experience is strongly preferred given our IBM Cloud footprint.



Reliability Engineering


  • Demonstrated use of SLIs, SLOs, and error budgets to drive decisions, with specific examples you can walk through.

  • Incident command experience on high‑severity, customer‑visible outages.



Automation and Infrastructure as Code


  • Strong Python and Bash for operational automation, at a standard you would trust to run unattended in production.

  • Terraform in a real team setting: modules, state management, peer review, and blast‑radius control. Ansible is a welcome addition.



Observability


  • Hands‑on with Prometheus and Grafana, plus at least one of ELK/EFK, OpenTelemetry, Datadog, CloudWatch, or

  • Ability to design alerting that engineers trust, and the judgement to delete alerts they do not.



Networking


  • Solid command of routing, DNS, TLS, load balancing, VPNs, and private interconnects, with the ability to debug a hybrid path methodically rather than by guesswork.



Security and Compliance


  • Comfort operating in a regulated environment, with HIPAA awareness or equivalent exposure to SOC 2, HITRUST, PCI DSS, or ISO 27001.



Ways of Working


  • Clear written communication across postmortems, design documents, and live incident updates.

  • Sound judgement under pressure, including the discipline to slow down when slowing down is the safer call.



Good‑to‑Have Skills


  • Red Hat OpenShift certification, or IBM Cloud Professional Architect or SRE certification.

  • AWS Professional or Specialty certifications.

  • Go for tooling, plus Kubernetes operator or controller development.

  • AIOps and automated remediation at scale, including anomaly detection on production telemetry.

  • Service mesh such as Istio or Linkerd, and progressive delivery with Argo Rollouts or Flagger.

  • GitOps practice with Argo CD or Flux.

  • FinOps discipline: cross‑cloud cost attribution, egress optimisation, and commitment planning.

  • Chaos engineering and structured game‑day programmes.

  • Experience migrating workloads between clouds, or repatriating them to on‑premises.

  • Background in healthcare, fintech, or another regulated domain.

  • Cloud: AWS (EKS, EC2, RDS, S3, VPC, Transit Gateway, Direct Connect) and IBM Cloud (Red Hat OpenShift on IBM

  • Containers and platform: Kubernetes, Red Hat OpenShift, Docker, Helm, Argo CD.

  • Infrastructure as code: Terraform (primary), Ansible, Git-based workflows.

  • Incident and delivery: PagerDuty, Opsgenie, Jira, Jenkins.

  • On‑premises: VMware vSphere, bare metal, hybrid DNS, and identity federation

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead DevOps Engineer
Lead DevOps Engineer

Lenskart • Gurugram District

On-site
INR 1,200,000 - 2,400,000
Engagement Manager - Support Lead
Engagement Manager - Support Lead

Quantiphi Analytics Solutions • Thiruvananthapuram

Hybrid
INR 2,500,000 - 3,500,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC Acceleration Centers • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Senior Site Reliability Engineer (SRE) / DevOps Engineer

Umanist Staffing LLC • Pune District

On-site
INR 2,250,000 - 2,750,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jibe Development Services • Navi Mumbai, Chennai District

On-site
INR 1,800,000 - 3,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior DevOps Engineer
Senior DevOps Engineer

Benchmarkit • Pune District

On-site
INR 1,400,000 - 2,400,000
Senior Staff Engineer, DevOps
Senior Staff Engineer, DevOps

Sierra Wireless • India

On-site
INR 2,500,000 - 4,000,000
Senior Cloud Infrastructure Engineer – Mumbai / Hybrid (3)
Senior Cloud Infrastructure Engineer – Mumbai / Hybrid (3)

Compoundexpress Private Limited • Mumbai

On-site
INR 1,400,000 - 2,500,000
Lead SRE / DevOps Architect
Lead SRE / DevOps Architect

V2 Solutions • Bengaluru

On-site
INR 4,000,000 - 7,000,000