Lead Site Reliability Engineer

FIS

Bengaluru

On-site

INR 350,000 - 700,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

FIS in Bengaluru, India, seeks a Senior Lead Site Reliability Engineer to define, build, and operate always-on, low-latency, and highly secure payment platforms powering large-scale financial transactions.

This senior role bridges distributed systems engineering, cloud platforms, and reliability architecture, setting technical direction, guiding incidents, and advancing platform maturity across Payments and FinTech.

Qualifications

  • Proven track record building and operating large-scale distributed systems.
  • Expertise in observability, alerting, and reliability engineering.
  • Cloud platforms expertise (AWS/Azure/GCP) and infrastructure-as-code.
  • Experience in Payments/FinTech or regulated environments.
  • Hands-on with Linux and Windows; complex enterprise stacks.
  • Leadership in incident management and post-incident reviews.
  • Ability to operate at staff level solving ambiguous problems across teams.

Responsibilities

  • Own and drive reliability outcomes at scale for real-time payment platforms.
  • Define reliability architecture and standards across services and infra.
  • Design and evolve observability platforms with actionable metrics and SLOs.
  • Lead and coordinate response to high-severity incidents and perform RCA.
  • Set strategy and drive adoption of SRE best practices and capacity modeling.
  • Architect automation and self-service platforms to enable frequent releases.
  • Partner with leaders to influence cloud migration and platform evolution.
  • Mentor senior engineers and raise overall reliability maturity.

Skills

Distributed systems
Observability
Cloud platforms
Linux
Incident management
Leadership

Tools

Prometheus
Grafana
Datadog
Splunk
ELK Stack
AWS
Azure
GCP
Terraform

Job description

About The Role

We are hiring a Senior Lead Site Reliability Engineer to define, build, and operate always-on, low-latency, and highly secure payment platforms that power large-scale financial transactions.

This is a senior technical role, not a pure operations position. You will operate at the intersection of distributed systems engineering, cloud platforms, and reliability architecture, setting technical direction and driving reliability outcomes across mission-critical, regulated systems in Payments and FinTech.

You will work across multiple teams and domains, influencing architecture, engineering practices, and operational maturity while remaining hands-on with the most complex reliability challenges.

What You Will Do
  • Own and drive reliability outcomes at scale for real-time, distributed payment and transaction processing platforms with strict SLAs, SLOs, and regulatory requirements.
  • Define reliability architecture and standards across services, platforms, and infrastructureshaping how systems are designed, deployed, observed, and operated.
  • Design and evolve enterprise-grade observability platforms (metrics, logs, traces, SLOs/SLIs) that provide actionable insights into system health, customer experience, and business impact.
  • Lead and coordinate response to high-severity production incidents, acting as a technical authority during major events and driving deep root-cause analysis and long-term systemic fixes.
  • Set strategy and drive adoption of SRE best practices including error budgets, capacity modeling, resilience testing, graceful degradation, and operational readiness.
  • Architect automation and self-service platforms that eliminate toil, reduce operational risk, and enable safe, frequent production releases across teams.
  • Partner with senior engineering, product, and platform leaders to influence architectural decisions, cloud migration strategy, disaster recovery posture, and long-term platform evolution.
  • Mentor senior engineers and technical leads, raising the overall reliability and operational maturity of the organization.
What You Bring
  • Deep software engineering expertise with a proven track record of building and operating large-scale, distributed, API-driven systems in production.
  • Expertise in observability, alerting, and reliability engineering, using tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or equivalent ecosystems.
  • Strong command of cloud platforms and open systems (AWS, Azure, or GCP), including infrastructure-as-code, platform automation, and cloud-native design patterns.
  • Significant experience running mission-critical systems in Payments, FinTech, Banking, or similarly regulated environments, where availability, correctness, and security are non-negotiable.
  • Hands-on experience across Linux (RHEL), Windows, databases (e.g., Oracle RDBMS), and complex enterprise stacks with strong system-level troubleshooting skills.
  • Demonstrated leadership in incident management, post-incident reviews, and continuous reliability improvement, with the ability to influence behavior and standards across teams.
  • Ability to operate effectively at Staff level scopesolving ambiguous problems, making trade-offs, and driving alignment across multiple teams and stakeholders.
Added Advantage
  • Strong automation and scripting skills using Python, Bash, Ansible, or similar tools.
  • Experience building or scaling CI/CD platforms and release automation in high-risk production environments.
  • Prior ownership of reliability strategy or platform initiatives spanning multiple teams or business units.
  • Experience modernizing legacy financial systems into cloud-native or hybrid architectures with a focus on resilience and compliance.
Why Join Us
  • Work on high-impact payment platforms operating at massive scale, where milliseconds and reliability directly affect real-world commerce.
  • Play a Staff-level role in defining reliability strategy for systems that cannot fail.
  • Join a culture that values engineering excellence, technical leadership, automation, and continuous learning.
Privacy Statement

FIS is committed to protecting the privacy and security of all personal information that we process in order to provide services to our clients. For specific information on how FIS protects personal information online, please see the Online Privacy Notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

fis • India

On-site
INR 4,000,000 - 6,500,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

FIS • Pune District

On-site
INR 3,500,000 - 6,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Impronics Technologies • Gurugram District

On-site
INR 1,500,000 - 2,500,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

SproutsAI • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

CellPoint Digital • Maharashtra

On-site
INR 3,000,000 - 6,000,000
Medical insurance with dependents
Head of DevOps & Site Reliability Engineering
Head of DevOps & Site Reliability Engineering

Finastra • Pune District

On-site
INR 3,500,000 - 6,500,000
Flexibility
Well‑being
Health & Financial Security
+4
Site Reliability Engineer -2
Site Reliability Engineer -2

Groww • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Senior Lead Java React Full Stack Engineer
Senior Lead Java React Full Stack Engineer

fis • Dadri

On-site
INR 4,200,000 - 7,000,000
GHMI/Hospitalization coverage for EMP-
Broad range of education and personal/
Lead Site Reliability Engineer SRE
Lead Site Reliability Engineer SRE

FIS • Chennai District

Hybrid
INR 3,500,000 - 5,500,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000