Site Reliability Engineer

Luxoft

Pune District

On-site

INR 3,500,000 - 5,500,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Luxoft seeks a Senior Site Reliability Engineer with 10+ years of experience to support a client in the Australian banking sector. The role concentrates on improving reliability, availability, and resilience through advanced monitoring and automated health checks.

You will design observability, implement IaC and CI/CD practices, and contribute to incident response, RCA, and AI-driven operations using modern tooling and cloud-native architectures.

Qualifications

  • SRE/Production Engineering experience with distributed systems, cloud-native tech, and Kubernetes/OpenShift.
  • Expertise in observability: metrics, logs, traces, dashboards, and anomaly detection using common tooling.
  • Knowledge of enterprise apps and middleware including FileNet, ICN, WAS, databases, and integration tech.
  • Ability to design for high availability, fault tolerance, and disaster recovery at scale.
  • Familiarity with IaC, DevSecOps, and Zero Trust security principles.
  • Proficient in CI/CD and reliability automation using Ansible and Chef.
  • Programming/scripting in Python, Java, Go, or PowerShell to support automation.
  • Strong troubleshooting skills across systems, apps, and networks.

Responsibilities

  • Improve service reliability and operational resilience via monitoring, resilience testing, DR planning, and SLO/SLI management.
  • Design observability capabilities: monitoring, alerting, dashboards, synthetic checks, and health visibility.
  • Develop automation and self-healing tooling; implement IaC and CI/CD to reduce manual effort.
  • Participate in major incident response, RCAs, and continuous improvement.
  • Lead AI-driven operational solutions using LLMs for triage and remediation.
  • Drive capacity, performance optimisation, and reliability improvements to meet demand.

Skills

SRE/Production Engineering
Distributed systems
Cloud-native
Kubernetes/OpenShift
Observability tooling
IaC / DevSecOps
CI/CD
Programming: Python/Java/Go/PowerShell
Troubleshooting

Tools

Elastic Stack
Splunk
Grafana
Prometheus
OpenTelemetry
Kubernetes
OpenShift
FileNet
ICN
WAS
CI/CD tooling
Ansible
Chef

Job description

Project description

We are seeking a Site Reliability Engineer with more than 10 years of experience to support our client in the Australian banking sector. The role centers on improving service reliability, availability, and operational resilience through comprehensive monitoring, resilience testing, disaster recovery planning, and management of SLIs and SLOs.

Responsibilities
  • 1. Improve service reliability, availability, and operational resilience through robust monitoring, resilience testing, disaster recovery planning, and management of SLIs, SLOs, and error budgets.
  • 2. Design and enhance observability capabilities, including monitoring, alerting, dashboards, synthetic monitoring, automated health checks, and visibility across customer journeys and cloud environments.
  • 3. Develop automation, self-healing capabilities, operational tooling, Infrastructure as Code and CI/CD practices to reduce manual effort and improve platform reliability.
  • 4. Participate in major incident response, service restoration, root cause analysis, and continuous improvement initiatives to minimise customer impact and improve recovery times.
  • 5. Design and implement AI-driven operational solutions, utilise LLMs for incident triage and remediation, and drive adoption of GitHub Copilot and AI-assisted engineering practices.
  • 6. Drive capacity management, performance optimisation and reliability improvements to ensure services meet current and future demand.
SKILLS
Must have
  • We are seeking a Site Reliability Engineer with more than 10 years of experience to support our client in the Australian banking sector.
  • 1. Demonstrated experience in Site Reliability Engineering (SRE), Production Engineering, DevOps, or Platform Engineering, with a strong understanding of distributed systems, cloud-native technologies, Kubernetes/OpenShift, and modern application architectures.
  • 2. Observability - Ability to use scripting and tooling to implement observability solutions, enabling the collection, analysis, and visualization of metrics, logs, and traces to support incident detection, diagnosis, and continuous service improvement. Familiarity with tools such as Elastic, Splunk, Grafana, Prometheus, OpenTelemetry, or similar enterprise monitoring platforms, and drive preventative improvements that enhance application performance and availability.
  • 3. Knowledge of enterprise application, middleware, and content management platforms, including FileNet, ICN, WAS, container platforms, databases and integration technologies.
  • 4. Reliability and Scalability - Ability to design and operate systems for high availability, fault tolerance, and disaster recovery, while ensuring systems can scale to meet current and future demand.
  • 5. Familiarity with Infrastructure as Code (IaC), DevSecOps practices and Zero Trust security principles.
  • 6. Proficiency in DevOps toolkits such as Ansible and Chef to support and drive CI/CD and reliability-automation initiatives.
  • 7. Programming and scripting languages including Python, Java, Go, PowerShell, or equivalent technologies to support automation and platform engineering initiatives.
  • 8. Troubleshooting - Capability to systematically identify, diagnose, and resolve technical issues across systems, applications, and networks, using analytical methods and tools to restore functionality, minimize disruption, and ensure stable operations.
Nice to have
  • Familiarity with Infrastructure as Code (IaC), DevSecOps practices and Zero Trust security principles.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Luxoft • Bengaluru

On-site
INR 900,000 - 1,300,000
Site Reliability Engineer
Site Reliability Engineer

Luxoft • Dadri

On-site
INR 2,500,000 - 4,200,000
Site Reliability Engineer
Site Reliability Engineer

Luxoft • Chennai District

On-site
INR 4,200,000 - 6,800,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Hyderabad

On-site
INR 1,400,000 - 2,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

Infosys • Hyderabad

On-site
INR 1,200,000 - 1,800,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Site Reliability Engineering Lead_Truist
Site Reliability Engineering Lead_Truist

Infosys • Bengaluru

On-site
INR 4,000,000 - 7,000,000