Senior Site Reliability Engineer (SRE) – Observability & Platform Systems

Saxo Group

Gurugram District

On-site

INR 3,500,000 - 5,500,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Saxo Bank is seeking a highly autonomous Senior Site Reliability Engineer in Gurugram, on-site. You will own our observability strategy, lead incident resolution, and drive reliability across core platform services in a regulated banking environment.

You will architect and optimize the observability stack, manage platforms like MongoDB, Kafka, Redis, Vault, and WSO2 within an OpenShift setup, and automate recovery using Python and Bash. Strong SRE/DevOps background and bank experience are valued.

Qualifications

  • 7+ years of experience in SRE, DevOps, or Platform Engineering.
  • Proven expertise in Grafana and the Prometheus ecosystem.
  • Experience defining and operating SLOs, SLIs, Error Budgets, and availability targets for production services.
  • Deep operational experience with Statefulset applications in production Kubernetes/OpenShift environments.
  • Proven track record of owning the full Incident Management lifecycle, including Root Cause Analysis and implementation of preventive measures.
  • Strong proficiency in Python and Bash for automation and tooling development.
  • Exceptional analytical skills with the ability to work independently and solve ambiguous technical challenges.

Responsibilities

  • Observability Leadership: Architect, maintain, and optimize end-to-end observability platform using Grafana, Prometheus, Elasticsearch.
  • Platform Engineering: Operate and engineer Redis, MongoDB, Kafka, Vault, and WSO2 within OpenShift environment for stability and security.
  • Incident Command: Lead L3 support and Major Incident Management from detection to post-mortem.
  • Automation: Build self-healing systems and automated recovery workflows using Python and Bash.
  • Proactive Reliability: Identify potential failure modes and implement architectural improvements to prevent outages.
  • Reliability Engineering: Define and track SLIs/SLOs and drive improvements with error budgets.

Skills

Grafana/Prometheus
Kubernetes/OpenShift
Python
Bash
Incident Management
Observability
Automation
Banking/Regulated industry

Tools

MongoDB
Kafka
Redis
HashiCorp Vault
WSO2 API Manager
OpenShift

Job description

Location - Gurugram (On-site) We are seeking a highly autonomous and technically profound Senior Site Reliability Engineer to own our observability strategy and critical platform services. In this role, you will not just monitor systems; you will engineer their reliability, performance, and security. You will serve as the primary authority for our monitoring stack and the operational lead for our core middleware layers. This position demands a proactive problem-solver who thrives in complex, regulated environments. You must be comfortable working independently, making high-stakes technical decisions, and leading incident resolution efforts from detection to post-mortem.

Responsibilities
  • Observability Leadership: Architect, maintain, and optimize our end-to-end observability platform using Grafana, Prometheus and Elasticsearch. You will define the standards for logging, metrics, and tracing.
  • Platform Engineering: Operate and engineer critical surrounding systems including MongoDB, Kafka, Redis, HashiCorp Vault, and WSO2 API Manager. You are responsible for their stability, scaling, and security within our OpenShift environment.
  • Incident Command: Lead L3 support and Major Incident Management. You will diagnose complex cross-layer issues (application to infrastructure) and drive permanent resolutions.
  • Automation: Build self-healing systems and automated recovery workflows using Python and Bash to reduce manual toil and improve system resilience.
  • Proactive Reliability: Identify potential failure modes before they occur and implement architectural improvements to prevent outages.
  • Reliability Engineering: Define and track service reliability objectives (SLIs/SLOs) and help engineering teams improve platform reliability through measurable outcomes and error budget management.
Your Profile
  • 7+ years of experience in SRE, DevOps, or Platform Engineering.
  • Proven expertise in Grafana and the Prometheus ecosystem.
  • Experience defining and operating SLOs, SLIs, Error Budgets, and availability targets for production services.
  • Deep operational experience with Statefulset applications in production Kubernetes/OpenShift environments.
  • Proven track record of owning the full Incident Management lifecycle, including Root Cause Analysis and implementation of preventive measures.
  • Strong proficiency in Python and Bash for automation and tooling development.
  • Exceptional analytical skills with the ability to work independently and solve ambiguous technical challenges without constant supervision.
  • Experience in the banking or highly regulated sectors is a strong plus.
What We Offer

The opportunity to work on a critical, large-scale digital banking platform. A culture that values engineering excellence, ownership, and proactive innovation. A collaborative international team environment with a focus on high-quality delivery. We get curious people invested in the world When you work at Saxo, you become a Saxonian and part of a purpose-driven organisation, where good ideas are always taken seriously, and where you can make a true impact. We are invested in your development, and you can expect a robust career from day one when you join Saxo – no matter which role you take on. You will join 2,500 other ambitious colleagues across 11 countries and become part of an international organisation. Working in Saxo, you will get to meet colleagues from many different cultures and backgrounds, and you should know that we value diversity and inclusion and see it as a genuine source of strength to drive growth, foster innovation and position us for long-term success. We encourage an open feedback culture and supportive team environments enabling employees to grow and fulfil their career aspirations. When you bring passion, curiosity, drive and team spirit, your learning journey will be dynamic and your career opportunities in Saxo will be immense. At Saxo we don’t just offer a job – we offer an opportunity to invest in your future!

About Saxo Bank

Support our purpose We believe access to global capital markets is not only for the privileged few. Our vision is to enable people to fulfill their financial aspirations to make an impact. That’s why we use the power of technology to deliver clients what they need, when they need it in a user-friendly and personalised experience. We aim to deliver the world’s most user-friendly and personalized trading and investment platform experience, which gives our clients exactly what they need to make more informed investment decisions.

Do you want to work for an extremely ambitious organisation at the cutting edge of banking and technology? You’ll be joining a company that:

  • Continuously strives to improve the Saxo Experience and exceed our clients’ expectations.
  • Always invests in the future.
  • Builds unrivalled platforms for seasoned and experienced investors and wholesale clients that provide real-time access to global capital markets.
  • Consistently wins the highest accolades for our platforms, products and services.

Own your future So, if you’re up for the challenge of breaking down barriers in global financial markets, we’re always looking for ambitious and enthusiastic people who share our passion for clients, technology and innovation and can bring new perspectives to our diverse team. For more on our story, culture and identity, please read the Saxo Bank Foundation, written by our CEO and founder Kim Fournais. More about Saxo Bank

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE) – Observability & Platform Systems
Senior Site Reliability Engineer (SRE) – Observability & Platform Systems

Saxo Group • India

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer (SRE) – Observability & Platform Systems
Senior Site Reliability Engineer (SRE) – Observability & Platform Systems

Saxo Bank • Gurugram District

On-site
INR 3,500,000 - 6,500,000
GCC Monitoring Analyst
GCC Monitoring Analyst

Saxo Group • Gurugram District

On-site
INR 1,200,000 - 2,100,000
Senior Client Service Associate
Senior Client Service Associate

Saxo Group • Gurugram District

On-site
INR 600,000 - 900,000
Senior Client Service Associate
Senior Client Service Associate

Saxo Bank • Gurugram District

On-site
INR 600,000 - 1,200,000
Corporate ODD & KYC
Corporate ODD & KYC

Saxo Group • Gurugram District

On-site
INR 600,000 - 1,200,000
Post Trade senior Analyst
Post Trade senior Analyst

Saxo Group India Private Limited (SGIPL) • Gurugram District

On-site
INR 600,000 - 900,000
Senior Post Trade Analyst
Senior Post Trade Analyst

Saxo Group • Gurugram District

On-site
INR 900,000 - 1,500,000
Senior Post Trade Analyst
Senior Post Trade Analyst

Saxo Bank • Gurugram District

On-site
INR 1,600,000 - 2,400,000
GCC Monitoring Analyst
GCC Monitoring Analyst

saxobank • Gurugram District

On-site
INR 1,500,000 - 2,100,000