Site Reliability Engineer (Azure Preferred)

FIS Solutions (India) Private Limited - Pune

Pune District

On-site

INR 2,500,000 - 6,000,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

FIS Solutions (India) Private Limited - Pune is seeking a Site Reliability Engineer (Azure Preferred) to ensure reliability and scalability of mission-critical banking, payments, and capital markets platforms. You will drive automation, strengthen operational resilience, and improve observability across cloud-native systems.

You will design monitoring, build automation, manage IaC, and lead incident management with a focus on high availability, performance, and security.

Qualifications

  • 5+ years of experience in Site Reliability Engineering, Production Support, Platform Engineering, DevOps, Cloud Operations, or a related field.
  • Hands-on experience with cloud platforms including AWS, Microsoft Azure, or Google Cloud Platform.
  • Strong knowledge of Infrastructure as Code and automation tools such as Terraform and Ansible.
  • Experience supporting web applications, APIs, distributed systems, and modern software architectures.
  • Proficiency with monitoring and observability tools including Prometheus, Grafana, Datadog, or similar platforms.
  • Experience with logging and analytics solutions such as Splunk, ELK Stack, or equivalent technologies.
  • Strong scripting and automation skills using Python, Bash, or similar programming languages.
  • Experience designing, maintaining, and optimizing CI/CD pipelines using Jenkins, GitLab CI/CD, Azure DevOps, or related tools.
  • Knowledge of containerization technologies such as Docker and container orchestration platforms.
  • Demonstrated experience in incident management, root cause analysis, and production support in enterprise environments.

Responsibilities

  • Design, implement, and maintain monitoring and observability solutions for infrastructure, applications, and customer experience.
  • Build and enhance automation frameworks to improve operational efficiency and reduce manual processes.
  • Ensure high availability, reliability, scalability, and performance of critical production systems.
  • Lead incident management activities, including triage, root cause analysis, recovery, and post-incident reviews.
  • Perform capacity planning, performance tuning, and infrastructure optimization to support business growth.
  • Develop and manage Infrastructure as Code solutions for consistent and scalable cloud deployments.
  • Maintain and optimize CI/CD pipelines to enable reliable and secure software delivery.
  • Collaborate with security teams to implement platform security controls and compliance best practices.
  • Develop, validate, and improve disaster recovery, backup, and business continuity strategies.
  • Partner with engineering, DevOps, QA, and product teams to achieve service-level objectives and operational excellence.
  • Participate in on-call rotations and provide support for critical production environments.

Skills

SRE experience
Cloud platforms
IaC & automation
Monitoring & observability
CI/CD pipelines
Containerization
Scripting: Python/Bash

Tools

Terraform
Ansible
Prometheus
Grafana
Datadog
Splunk
ELK Stack
Jenkins
GitLab CI/CD
Azure DevOps
Docker
Kubernetes

Job description

Site Reliability Engineer (Azure Preferred) – 5+ Yrs – Pune Location

About the Role As a Site Reliability Engineer, you will be responsible for ensuring the reliability, scalability, performance, and availability of mission-critical banking, payments, and capital markets platforms. You will drive automation, strengthen operational resilience, and improve observability across cloud-native and distributed systems. Working closely with engineering, DevOps, security, QA, and product teams, you will help deliver highly available services while reducing operational risk. Success in this role is measured through platform stability, service reliability, incident reduction, and continuous operational improvement.

What You Will Be Doing
  • Design, implement, and maintain monitoring and observability solutions for infrastructure, applications, and customer experience.
  • Build and enhance automation frameworks to improve operational efficiency and reduce manual processes.
  • Ensure high availability, reliability, scalability, and performance of critical production systems.
  • Lead incident management activities, including triage, root cause analysis, recovery, and post-incident reviews.
  • Perform capacity planning, performance tuning, and infrastructure optimization to support business growth.
  • Develop and manage Infrastructure as Code solutions for consistent and scalable cloud deployments.
  • Maintain and optimize CI/CD pipelines to enable reliable and secure software delivery.
  • Collaborate with security teams to implement platform security controls and compliance best practices.
  • Develop, validate, and improve disaster recovery, backup, and business continuity strategies.
  • Partner with engineering, DevOps, QA, and product teams to achieve service-level objectives and operational excellence.
  • Participate in on-call rotations and provide support for critical production environments.
What you bring
  • 5+ years of experience in Site Reliability Engineering, Production Support, Platform Engineering, DevOps, Cloud Operations, or a related field.
  • Hands-on experience with cloud platforms including AWS, Microsoft Azure, or Google Cloud Platform.
  • Strong knowledge of Infrastructure as Code and automation tools such as Terraform and Ansible.
  • Experience supporting web applications, APIs, distributed systems, and modern software architectures.
  • Proficiency with monitoring and observability tools including Prometheus, Grafana, Datadog, or similar platforms.
  • Experience with logging and analytics solutions such as Splunk, ELK Stack, or equivalent technologies.
  • Strong scripting and automation skills using Python, Bash, or similar programming languages.
  • Experience designing, maintaining, and optimizing CI/CD pipelines using Jenkins, GitLab CI/CD, Azure DevOps, or related tools.
  • Knowledge of containerization technologies such as Docker and container orchestration platforms.
  • Demonstrated experience in incident management, root cause analysis, and production support in enterprise environments.
Preferred Qualifications
  • Experience in applying SRE practices, reliability engineering principles, and service-level management in large-scale environments.
  • Knowledge of Kubernetes, cloud-native architectures, and microservices-based platforms.
  • Experience conducting operational readiness assessments and post-mortem reviews.
  • Understanding of disaster recovery, high-availability design patterns, and resiliency engineering.
  • Industry certifications in AWS, Azure, Google Cloud, Kubernetes, DevOps, or Site Reliability Engineering
What we offer you

A work environment built on collaboration, flexibility and respect Competitive salary and attractive range of benefits designed to help support your lifestyle and wellbeing Varied and challenging work to help you grow your technical skillset

Privacy Statement

FIS is committed to protecting the privacy and security of all personal information that we process in order to provide services to our clients. For specific information on how FIS protects personal information online, please see the Online Privacy Notice.

Sourcing Model

Recruitment at FIS works primarily on a direct sourcing model; a relatively small portion of our hiring is through recruitment agencies. FIS does not accept resumes from recruitment agencies which are not on the preferred supplier list and is not responsible for any related fees for resumes submitted to job postings, our employees, or any other part of our company.

#pridepass

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (Azure Preferred)
Site Reliability Engineer (Azure Preferred)

fis • Pune District

On-site
INR 1,200,000 - 1,800,000
Competitive salary
Flexible work environment
Career growth opportunities
Site Reliability Engineer
Site Reliability Engineer

fis • Pune District

On-site
INR 1,200,000 - 1,800,000
Private medical cover
Dental cover
Travel insurance
+1
Site Reliability Engineer (Java,Unix,Dynatrace and Splunk)
Site Reliability Engineer (Java,Unix,Dynatrace and Splunk)

FIS • Pune District

On-site
INR 1,000,000 - 1,400,000
Opportunity in a leading FinTech company
Professional development possibilities
High degree of responsibility
MS Azure Operational Analyst – 24/7 Rotational Shifts- Pune
MS Azure Operational Analyst – 24/7 Rotational Shifts- Pune

FIS Solutions (India) Private Limited - Pune • Pune District

On-site
INR 1,200,000 - 1,800,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

FIS • Pune District

On-site
INR 3,500,000 - 6,000,000
MS Azure Operational Analyst - 24/7 Rotational Shifts- Pune
MS Azure Operational Analyst - 24/7 Rotational Shifts- Pune

fis • Pune District

On-site
INR 900,000 - 1,200,000
Site Reliability Engineer – Windows
Site Reliability Engineer – Windows

UBS • Maharashtra

On-site
INR 2,500,000 - 4,000,000
Senior Enterprise Platform Support Engineer – AI & Cloud
Senior Enterprise Platform Support Engineer – AI & Cloud

FIS • Bengaluru

On-site
INR 3,000,000 - 5,200,000
Competitive salary
Professional learning
Inclusive, diverse environment
+2
MS Azure Operational Analyst – 24/7 Rotational Shifts- Pune
MS Azure Operational Analyst – 24/7 Rotational Shifts- Pune

FIS • Pune District

On-site
INR 900,000 - 1,300,000
Senior Enterprise Platform Support Engineer – AI & Cloud
Senior Enterprise Platform Support Engineer – AI & Cloud

FIS Solutions (India) Private Limited - Bangalore • Pune District

On-site
INR 4,000,000 - 8,000,000