Senior Site Reliability Engineer

BayOne Solutions

Hyderabad

On-site

INR 3,000,000 - 6,000,000

Full time

21 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

BayOne Solutions seeks a Senior Site Reliability Engineer to lead reliability for core banking infrastructure and distributed transaction processing applications.

You will design self-healing architectures, define SLOs/SLAs, and drive zero-downtime operations, with emphasis on incident leadership and automated compliance. The role demands 5+ years in SRE/DevOps in fintech, deep Kubernetes and cloud experience, and hands-on IaC with Terraform/Ansible.

Qualifications

  • 5+ years in SRE/DevOps/Engineering.
  • Experience in banking/fintech or high-volume transactional environments.
  • Kubernetes and cloud architectures (AWS/Azure/GCP).
  • Programming in Go, Python, or Java for tooling.
  • Understanding of relational and distributed databases (PostgreSQL, Cassandra, Redis).
  • Observability with ELK/EFK, OpenTelemetry, Prometheus, and tracing.

Responsibilities

  • Design fault-tolerant, multi-region distributed banking systems.
  • Define and manage SLOs/SLIs and error budgets with product managers.
  • Lead incident response for P1/P2 outages and drive blameless postmortems.
  • Implement IaC and automated self-healing mechanisms.
  • Forecast capacity, run load tests, and perform chaos engineering experiments.
  • Ensure compliance with PCI-DSS, SOC2 and regulatory guidelines.

Skills

Kubernetes
SRE/DevOps
IaC
Public Cloud
Observability
Chaos engineering
Go/Python/Java

Education

Bachelor/Master in CS/Engineering

Tools

Terraform
Ansible
Prometheus
OpenTelemetry
Jaeger
ELK Stack
Kafka
PostgreSQL
Redis

Job description

We are seeking a Senior Site Reliability Engineer with 5+ years of experience to lead the reliability engineering strategy for our core banking infrastructure and distributed transaction processing applications. You will design self-healing architectures, establish SLOs/SLAs, lead major incident resolution, and drive zero-downtime architecture for mission-critical financial platforms.

Key Responsibilities
  • Reliability Architecture & Design: Partner with software architects to design fault-tolerant, multi-region distributed banking systems capable of processing millions of daily transactions.
  • SLO & Error Budget Management: Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets alongside product managers to balance feature velocity with system stability.
  • Advanced Automation & IaC: Implement Infrastructure as Code (IaC) using Terraform, Ansible, and Kubernetes operators. Build automated self-healing mechanisms to resolve known failure modes without human intervention.
  • Incident Leadership & Postmortems: Lead response for high-priority (P1/P2) banking outages. Facilitate blameless post-mortems and enforce long-term root cause remediations.
  • Capacity & Chaos Engineering: Forecast system growth, conduct load testing under peak banking hours, and run Chaos Engineering experiments (Gremlin, Chaos Mesh) to uncover hidden vulnerabilities.
  • Compliance & Security: Ensure infrastructure complies with banking regulatory frameworks (PCI-DSS, SOC2, Central Bank guidelines) and lead automated compliance auditing tools.
Required Qualifications & Skills
  • Education: Btech or Bachelor’s or Master’s degree in Computer Science, Engineering, or equivalent practical experience.
  • Deep Experience: 5+ years in SRE, DevOps, or Software Engineering, with at least 2 years in banking, fintech, or high-volume transactional environments.
  • Orchestration & Cloud: Expert-level knowledge of Kubernetes (CKA certified preferred) and public/hybrid cloud enterprise architectures (AWS/Azure/GCP).
  • Software Development: Strong programming skills in Go, Python, or Java for building internal SRE tooling, CLI utilities, and automated controllers.
  • Data Stores: Understanding of high-availability relational (PostgreSQL, Oracle) and distributed non-relational databases (Cassandra, Redis, Kafka).
  • Observability: Mastery of ELK/EFK stack, OpenTelemetry, Prometheus, Cortex/Thanos, and distributed tracing (Jaeger/Zipkin).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Impronics Technologies • Gurugram District

On-site
INR 1,500,000 - 2,500,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Delivery Lead-SRE
Delivery Lead-SRE

Acuity Analytics • Bengaluru

On-site
INR 2,400,000 - 3,800,000
Site Reliability Engineer – Windows
Site Reliability Engineer – Windows

UBS • Maharashtra

On-site
INR 2,500,000 - 4,000,000
Resilience and Reliability Engineer
Resilience and Reliability Engineer

EY • Pune District, Gurugram District, Bengaluru

Hybrid
INR 1,800,000 - 2,800,000
Site Reliability Engineer - Vice President
Site Reliability Engineer - Vice President

Citi • Maharashtra

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Luxoft • Pune District

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Luxoft • Bengaluru

On-site
INR 900,000 - 1,300,000