Analyst II, Production Support

FIS Solutions (India) Private Limited - Pune

Pune District

On-site

INR 1,200,000 - 2,400,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

FIS Solutions (India) Private Limited - Pune is seeking a Site Reliability & Cloud Engineer to design, deploy, and operate reliable cloud services. The role emphasizes automation, observability and secure platform engineering across AWS and/or Azure.

Responsibilities include defining SLIs/SLOs, building runbooks, and supporting incident response with a focus on availability and performance. Collaboration with security, development, and operations teams is essential.

Qualifications

  • Three or more years in cloud/production engineering, DevOps or SRE.
  • Experience operating production workloads on AWS and/or Azure.
  • Strong Linux/UNIX administration and troubleshooting.

Responsibilities

  • Own reliability, availability, performance and readiness of cloud-hosted apps and platform services.
  • Define and maintain SLIs/SLOs, availability targets and health dashboards.
  • Build monitoring, alerting and diagnostics to reduce toil and improve incident response.
  • Automate deployments, configurations and evidence collection using code and pipelines.
  • Design and maintain cloud infrastructure with IaC and secure engineering standards.
  • Participate in on‑call rotation and incident response with clear stakeholder communications.

Skills

Cloud engineering
AWS
Azure
CI/CD
Python
Linux
High availability
Observability
Incident management

Tools

Kubernetes (EKS/AKS)
Docker
Terraform

Job description

SITE RELIABILITY & CLOUD ENGINEER Production Engineering | Pune, India

Engineer reliable, secure and scalable cloud services by combining production operations with software engineering, automation, observability and cloud platform expertise. The role owns service reliability outcomes across the application lifecycle and partners with development, infrastructure, security and support teams to prevent incidents, reduce toil and improve customer experience.

Role Overview

We are seeking an experienced Site Reliability and Cloud Engineer with strong hands‑on expertise in AWS and/or Microsoft Azure. The successful candidate will design, deploy, operate and continuously improve production services, with a focus on availability, performance, resilience, security and operational efficiency. This is an engineering‑led production role. The engineer will use automation, infrastructure as code, telemetry and disciplined incident/problem management to create sustainable services, accelerate safe change and reduce manual operational effort.

Key Responsibilities
  • Own the reliability, availability, performance and operational readiness of cloud-hosted applications and platform services.
  • Define and maintain service level indicators (SLIs), service level objectives (SLOs), availability targets and actionable service-health dashboards.
  • Build monitoring and alerting around customer‑impacting symptoms, golden signals and service dependencies; reduce alert noise and improve diagnostic quality.
  • Automate repeatable operational work, remediation, deployments, configuration, evidence collection and environment validation using code and pipelines.
  • Design, build and maintain cloud infrastructure using infrastructure as code, reusable modules, policy guardrails and secure engineering standards.
  • Participate in incident response and on‑call support, including rapid triage, stabilisation, technical escalation and clear stakeholder communication.
  • Lead or contribute to blameless post‑incident reviews; identify root causes, track corrective actions and engineer controls that prevent recurrence.
  • Partner with application engineering teams on architecture, capacity planning, performance engineering, release readiness and production operability.
  • Improve deployment safety through CI/CD controls, automated testing, progressive validation, rollback strategies and change‑risk reduction.
  • Engineer resilience through high‑availability patterns, backup and restore validation, disaster‑recovery runbooks, failover testing and dependency mapping.
  • Manage production risks including vulnerabilities, patching, certificates, secrets, access controls, audit evidence and cloud security findings.
  • Create and maintain runbooks, standard operating procedures, architecture records and operational knowledge that support consistent 24x7 service delivery.
  • Analyse operational data and trends to reduce mean time to detect and restore service, eliminate recurring failure modes and improve capacity and cost efficiency.
  • Mentor support and engineering colleagues in SRE practices, automation, observability, troubleshooting and operational ownership.
Required Skills and Experience
  • Three or more years of experience in cloud engineering, production engineering, DevOps, platform engineering or Site Reliability Engineering.
  • Hands‑on experience operating production workloads on AWS and/or Microsoft Azure, including compute, networking, identity, storage, managed databases and monitoring services.
  • Strong Linux/UNIX administration and troubleshooting skills; working knowledge of Windows Server is beneficial.
  • Practical experience with Kubernetes and containers, including EKS and/or AKS, Docker, deployment troubleshooting and workload reliability.
  • Infrastructure‑as‑code experience using Terraform; ability to build reusable, governed and maintainable modules.
  • CI/CD experience with tools such as Harness, Azure DevOps, GitHub Actions, Jenkins or equivalent, including deployment and rollback controls.
  • Programming or advanced automation capability using Python, PowerShell, Bash or a comparable language; coding experience beyond simple one‑off scripts.
  • Experience with observability platforms and practices covering metrics, logs, traces, dashboards, alerting and application performance monitoring.
  • Strong incident and problem management experience, including technical triage, root cause analysis, corrective actions and production communications.
  • Understanding of distributed systems, scalability, high availability, capacity management, performance bottlenecks and failure modes.
  • Experience supporting SQL Server and cloud database services, including connectivity, performance diagnostics, backup/restore and operational monitoring.
  • Working knowledge of cloud security, least privilege, certificate and secrets management, vulnerability remediation, auditing and compliance controls.
  • Experience with ServiceNow or a comparable IT service management platform for incidents, problems, changes and operational work tracking.
  • Clear written and verbal communication, disciplined documentation, strong ownership and the ability to work across engineering and business teams.
Desirable Skills
  • Cloud certification in AWS or Microsoft Azure; Kubernetes or Terraform certification is advantageous.
  • Experience with Akamai, API gateways, web application delivery, DNS, load balancing, VPNs and enterprise network connectivity.
  • Knowledge of event‑driven and service‑oriented architectures, domain‑driven design and messaging platforms.
  • Experience in regulated financial services or another environment with formal change, risk, audit, resilience and data‑protection obligations.
  • Experience implementing policy as code, security scanning, automated compliance controls, FinOps or cloud cost optimisation.
  • Experience supporting globally distributed services, customer onboarding and follow‑the‑sun operational models.
Success Measures
  • Outcome Evidence of Success Reliability Services have defined health measures, meaningful alerts, tested recovery procedures and improving availability trends.
  • Engineering efficiency Manual toil and recurring incidents are reduced through automation, reusable tooling and preventive engineering.
  • Operational readiness Releases and services meet documented production‑readiness, security, observability and support requirements.
  • Incident learning Major incidents produce clear root causes, owned corrective actions and measurable risk reduction.
  • Collaboration Development, platform, security and support teams share clear ownership and use consistent operational practices.
Role Expectations
  • Location: Pune, India.
  • Flexible to work on rotational shifts and weekends too.
  • Work closely with global engineering, production support, security and service‑management teams.
  • Participate in an agreed on‑call or out‑of‑hours support rotation where required for production services.
  • Demonstrate an engineering mindset: automate where practical, design for failure, measure service health and treat operational learning as product improvement.
Privacy Statement

FIS is committed to protecting the privacy and security of all personal information that we process in order to provide services to our clients. For specific information on how FIS protects personal information online, please see the Online Privacy Notice.

Sourcing Model

Recruitment at FIS works primarily on a direct sourcing model; a relatively small portion of our hiring is through recruitment agencies. FIS does not accept resumes from recruitment agencies which are not on the preferred supplier list and is not responsible for any related fees for resumes submitted to job postings, our employees, or any other part of our company.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Analyst II, Production Support
Analyst II, Production Support

fis • Pune District

On-site
INR 1,500,000 - 2,300,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

fis • India

On-site
INR 4,000,000 - 6,500,000
MS Azure Operational Analyst - 24/7 Rotational Shifts- Pune
MS Azure Operational Analyst - 24/7 Rotational Shifts- Pune

fis • Pune District

On-site
INR 900,000 - 1,200,000
MS Azure Operational Analyst – 24/7 Rotational Shifts- Pune
MS Azure Operational Analyst – 24/7 Rotational Shifts- Pune

FIS Solutions (India) Private Limited - Pune • Pune District

On-site
INR 1,200,000 - 1,800,000
Production Support - SQL and Oracle PL/SQ (5 AM to 2 PM) Shit - Pune
Production Support - SQL and Oracle PL/SQ (5 AM to 2 PM) Shit - Pune

fis • Pune District

On-site
INR 900,000 - 1,300,000
Growth potential
Professional development
Competitive salary & benefits
+1
Senior Enterprise Platform Support Engineer – AI & Cloud
Senior Enterprise Platform Support Engineer – AI & Cloud

FIS • Bengaluru

On-site
INR 3,000,000 - 5,200,000
Competitive salary
Professional learning
Inclusive, diverse environment
+2
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

FIS • Chennai District

Hybrid
INR 3,500,000 - 6,500,000
Collaboration culture
Competitive salary & benefits
Career growth opportunities
Cloud Engineer (AWS)
Cloud Engineer (AWS)

FIS • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Production Support Supervisor I
Production Support Supervisor I

FIS • Maharashtra

On-site
INR 800,000 - 1,200,000
Senior Enterprise Platform Support Engineer AI & Cloud
Senior Enterprise Platform Support Engineer AI & Cloud

FIS • Bengaluru

On-site
INR 3,000,000 - 4,500,000