Analyst II, Production Support

fis

Pune District

On-site

INR 1,500,000 - 2,300,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fis is seeking an experienced Site Reliability & Cloud Engineer in Pune to design, deploy and operate scalable cloud services with a focus on availability, security and performance. You will automate infrastructure, implement observability, and contribute to incident response and post-incident reviews.

The role requires hands-on AWS and/or Azure expertise, Kubernetes/Docker, Terraform and a strong automation mindset to reduce toil and improve customer experience.

Qualifications

  • 3+ years in cloud engineering, production engineering or SRE.
  • Hands-on with AWS and/or Azure; strong Linux/UNIX.
  • Experience with Kubernetes, Docker and CI/CD tooling.
  • Infrastructure as code with Terraform.

Responsibilities

  • Own reliability and availability of cloud-hosted apps and services.
  • Define SLIs/SLOs and health dashboards.
  • Automate operations, deployments and validation with code.
  • Design and maintain cloud infrastructure using IaC and guardrails.
  • Participate in incident response and on-call support.
  • Lead or contribute to blameless post-incident reviews.
  • Collaborate with app teams on performance and operability.
  • Improve deployment safety with CI/CD controls and testing.

Skills

Cloud engineering
DevOps
SRE
Platform engineering
Incident management

Tools

Terraform
Kubernetes
Docker
AWS
Azure
GitHub Actions
Jenkins

Job description

SITE RELIABILITY & CLOUD ENGINEER

Production Engineering | Pune, India

ROLE PURPOSE

Engineer reliable, secure and scalable cloud services by combining production operations with software engineering, automation, observability and cloud platform expertise. The role owns service reliability outcomes across the application lifecycle and partners with development, infrastructure, security and support teams to prevent incidents, reduce toil and improve customer experience.

Role Overview

We are seeking an experienced Site Reliability and Cloud Engineer with strong hands‑on expertise in AWS and/or Microsoft Azure. The successful candidate will design, deploy, operate and continuously improve production services, with a focus on availability, performance, resilience, security and operational efficiency. This is an engineering‑led production role. The engineer will use automation, infrastructure as code, telemetry and disciplined incident/problem management to create sustainable services, accelerate safe change and reduce manual operational effort.

Key Responsibilities
  • Own the reliability, availability, performance and operational readiness of cloud-hosted applications and platform services.
  • Define and maintain service level indicators (SLIs), service level objectives (SLOs), availability targets and actionable service-health dashboards.
  • Build monitoring and alerting around customer-impacting symptoms, golden signals and service dependencies; reduce alert noise and improve diagnostic quality.
  • Automate repeatable operational work, remediation, deployments, configuration, evidence collection and environment validation using code and pipelines.
  • Design, build and maintain cloud infrastructure using infrastructure as code, reusable modules, policy guardrails and secure engineering standards.
  • Participate in incident response and on‑call support, including rapid triage, stabilisation, technical escalation and clear stakeholder communication.
  • Lead or contribute to blameless post‑incident reviews; identify root causes, track corrective actions and engineer controls that prevent recurrence.
  • Partner with application engineering teams on architecture, capacity planning, performance engineering, release readiness and production operability.
  • Improve deployment safety through CI/CD controls, automated testing, progressive validation, rollback strategies and change‑risk reduction.
  • Engineer resilience through high‑availability patterns, backup and restore validation, disaster‑recovery runbooks, failover testing and dependency mapping.
  • Manage production risks including vulnerabilities, patching, certificates, secrets, access controls, audit evidence and cloud security findings.
  • Create and maintain runbooks, standard operating procedures, architecture records and operational knowledge that support consistent 24x7 service delivery.
  • Analyse operational data and trends to reduce mean time to detect and restore service, eliminate recurring failure modes and improve capacity and cost efficiency.
  • Mentor support and engineering colleagues in SRE practices, automation, observability, troubleshooting and operational ownership.
Required Skills and Experience
  • Three or more years of experience in cloud engineering, production engineering, DevOps, platform engineering or Site Reliability Engineering.
  • Hands‑on experience operating production workloads on AWS and/or Microsoft Azure, including compute, networking, identity, storage, managed databases and monitoring services.
  • Strong Linux/UNIX administration and troubleshooting skills; working knowledge of Windows Server is beneficial.
  • Practical experience with Kubernetes and containers, including EKS and/or AKS, Docker, deployment troubleshooting and workload reliability.
  • Infrastructure‑as‑code experience using Terraform; ability to build reusable, governed and maintainable modules.
  • CI/CD experience with tools such as Harness, Azure DevOps, GitHub Actions, Jenkins or equivalent, including deployment and rollback controls.
  • Programming or advanced automation capability using Python, PowerShell, Bash or a comparable language; coding experience beyond simple one‑off scripts.
  • Experience with observability platforms and practices covering metrics, logs, traces, dashboards, alerting and application performance monitoring.
  • Strong incident and problem management experience, including technical triage, root cause analysis, corrective actions and production communications.
  • Understanding of distributed systems, scalability, high availability, capacity management, performance bottlenecks and failure modes.
  • Experience supporting SQL Server and cloud database services, including connectivity, performance diagnostics, backup/restore and operational monitoring.
  • Working knowledge of cloud security, least privilege, certificate and secrets management, vulnerability remediation, auditing and compliance controls.
  • Experience with ServiceNow or a comparable IT service management platform for incidents, problems, changes a
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Analyst II, Production Support
Analyst II, Production Support

FIS Solutions (India) Private Limited - Pune • Pune District

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Senior Site Reliability Engineer (SRE) / DevOps Engineer

Umanist Staffing LLC • Pune District

On-site
INR 2,250,000 - 2,750,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Senior Site Reliability Engineer (SRE) / DevOps Engineer

Umanist Staffing LLC • Maharashtra

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

MishiPay • Bengaluru

On-site
INR 2,500,000 - 3,800,000
Site Reliability Engineer Lead (Immediate Joiner)
Site Reliability Engineer Lead (Immediate Joiner)

F-Prime Capital • Pune District

On-site
INR 3,500,000 - 6,000,000