Senior Site Reliability Engineer - Infrastructure

S&P Global Market Intelligence

Bengaluru, Hyderabad

On-site

INR 1,500,000 - 2,500,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

A global data analytics firm in Bengaluru is seeking a Senior Site Reliability Engineer. This role entails ensuring the reliability and scalability of production systems, working closely with Infrastructure, Application, and Security teams. Candidates should have over 6 years of SRE or DevOps experience, strong Python skills for automation, and expertise in AWS and Kubernetes. The position offers an exciting opportunity to drive operational excellence in a fast-paced environment.

Qualifications

  • 6+ years of experience in SRE, DevOps, Platform, or Infrastructure Engineering roles.
  • Strong software engineering background with a focus on Python.
  • Experience with scalable, distributed systems in production.

Responsibilities

  • Own and operate production services supporting critical applications.
  • Design, build, and manage AWS infrastructure, including EKS-based clusters.
  • Monitor system health and troubleshoot complex issues.

Skills

AWS Cloud environments
Python development
Kubernetes
Infrastructure automation
Incident management

Tools

Terraform
PostgreSQL
Prometheus
Grafana

Job description

About the Role

As a Senior Site Reliability Engineer (SRE) at Kensho, you will be a hands-on technologist who combines strong infrastructure expertise with solid software engineering skills Python first. You will be responsible for ensuring the reliability, scalability, and security of both business-critical internal systems and external, customer facing services.

You will work closely with Infrastructure, Application, and Security teams to design resilient systems, automate operations, and continuously improve platform stability. This role requires deep ownership of production systems, strong troubleshooting skills across infrastructure, Container orchestration systems, networking, and applications, and comfort operating in a 24/7 on call environment.

What You’ll Do
  • Own and operate production services supporting critical financial applications with a strong focus on availability, performance, and reliability
  • Design, build, and manage AWS infrastructure, including EKS-based clusters, across lower and production environments
  • Provision and manage infrastructure using Terraform (Infrastructure as Code) with a strong automation first mindset
  • Deploy, scale, and troubleshoot applications running on like Kubernetes, including cluster creation, upgrades, and lifecycle management
  • Build and maintain automation frameworks and tooling Python based to reduce operational toil and prevent recurring incidents
  • Monitor system health using metrics, logs, and alerts; continuously tune alerts, dashboards, and runbooks
  • Troubleshoot complex issues spanning clusters, networking, certificates, deployments, and application behavior
  • Manage certificate lifecycle and expiration, ensuring secure and uninterrupted service operation
  • Collaborate with InfoSec, Vulnerability Management, and Network Security teams (e.g., Zscaler) to maintain a strong security posture
  • Collaborate with L1/L2 teams, helping them understand infrastructure and operational best practices
  • Participate in on call and lead incident response, drive root cause analysis, and ensure effective post-incident remediation and learnings
  • Identify architectural anti-patterns and drive improvements by reviewing new services for production readiness, resiliency, and secure design prior to release
  • Establish and enforce production readiness standards, including deployment strategies, rollback plans, and observability requirements
  • Optimize infrastructure cost and resource utilization without compromising reliability and performance
What We Look For
  • 6+ years of experience in SRE, DevOps, Platform, or Infrastructure Engineering roles
  • Strong software engineering background, with hands-on Python development used for automation, tooling, and system reliability
  • Experience building or supporting scalable, distributed systems in production
  • Deep experience with AWS cloud environments, including AWS, IAM, networking, and access controls
  • Strong hands-on expertise with similar tools like Kubernetes (EKS preferred): cluster creation, deployments, scaling, and troubleshooting
  • Solid understanding of networking fundamentals (VPCs, routing, DNS, load balancing, security groups)
  • Experience with CI/CD pipelines, deployment tools, and infrastructure automation
  • Working knowledge of databases and query optimization, and understanding how applications behave under load
  • Familiarity with similar tools like Kafka or other messaging systems
  • Comfortable conducting code reviews and participating in coding focused interviews
  • Strong operational mindset with experience in incident management and oncall rotations
  • Clear communicator and collaborative teammate who values documentation and knowledge sharing
How to Really Get Our Attention
  • Demonstrated ownership of largescale, production systems
  • Strong examples of Python based automation or internal tooling
  • Contributions to opensource projects, infrastructure platforms, or reliability tooling
  • Experience working closely with security and compliance teams in regulated environments
Technologies We Like
  • AWS, Amazon EKS, Terraform, Jsonnet
  • Similar tools like Kubernetes, Helm, CI/CD tooling
  • Python (automation, tooling, reliability engineering)
  • Prometheus, Grafana, logging and monitoring platforms
  • PostgreSQL and other production databases
  • Similar tools like Kafka or event driven systems
  • Linux (Ubuntu or similar)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer- Site Reliability
Senior Software Engineer- Site Reliability

S&P Global • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Health care coverage
Generous time off
Access to career resources
+2
Senior Software Engineer- Site Reliability
Senior Software Engineer- Site Reliability

S&P Global • Bengaluru

On-site
INR 6,535,000 - 10,271,000
Health & Wellness
Flexible Downtime
Continuous Learning
+3
Site Reliability Engineer (SRE) / DevOps Engineer
Site Reliability Engineer (SRE) / DevOps Engineer

New Era Technology • Gurugram District

On-site
INR 1,500,000 - 2,100,000
AWS Site Reliability Engineer (SRE)
AWS Site Reliability Engineer (SRE)

Zensar • Hyderabad, Pune District

Hybrid
INR 1,500,000 - 2,100,000
Site Reliability Engineer
Site Reliability Engineer

Smart Ims • Bengaluru

Hybrid
INR 1,200,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Metlife • Hyderabad

Hybrid
INR 1,500,000 - 2,600,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Chennai District

On-site
INR 2,500,000 - 4,000,000
Senior Manager - Site Reliability Engineer|NR-2026-0246
Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net • Bengaluru

On-site
INR 6,000,000 - 8,000,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Okta • Bengaluru

On-site
INR 1,500,000 - 2,500,000