About The Role
Senior Site Reliability Engineer (SRE) at Kensho, a part of S&P Global, will be a hands-on technologist blending infrastructure expertise with solid software engineering skills (Python first). You will be responsible for the reliability, scalability, and security of both business-critical internal systems and external, customer facing services.
You will work closely with Infrastructure, Application, and Security teams to design resilient systems, automate operations, and continuously improve platform stability. This role requires deep ownership of production systems, strong troubleshooting skills across infrastructure, container orchestration systems, networking, and applications, and comfort operating in a 24/7 on call environment.
What You’ll Do
- Own and operate production services supporting critical financial applications with a strong focus on availability, performance, and reliability
- Design, build, and manage AWS infrastructure, including EKS-based clusters, across lower and production environments
- Provision and manage infrastructure using Terraform (Infrastructure as Code) with a strong automation first mindset
- Deploy, scale, and troubleshoot applications running on Kubernetes, including cluster creation, upgrades, and lifecycle management
- Build and maintain automation frameworks and tooling Python-based to reduce operational toil and prevent recurring incidents
- Monitor system health using metrics, logs, and alerts; continuously tune alerts, dashboards, and runbooks
- Troubleshoot complex issues spanning clusters, networking, certificates, deployments, and application behavior
- Manage certificate lifecycle and expiration, ensuring secure and uninterrupted service operation
- Collaborate with InfoSec, Vulnerability Management, and Network Security teams to maintain a strong security posture
- Collaborate with L1/L2 teams, helping them understand infrastructure and operational best practices
- Participate in on-call and lead incident response, drive root cause analysis, and ensure effective post-incident remediation and learnings
- Identify architectural anti-patterns and drive improvements by reviewing new services for production readiness, resiliency, and secure design prior to release
- Establish and enforce production readiness standards, including deployment strategies, rollback plans, and observability requirements
- Optimize infrastructure cost and resource utilization without compromising reliability and performance
What We Look For
- 6+ years of experience in SRE, DevOps, Platform, or Infrastructure Engineering roles
- Strong software engineering background, with hands-on Python development used for automation, tooling, and system reliability
- Experience building or supporting scalable, distributed systems in production
- Deep experience with AWS cloud environments, including IAM, networking, and access controls
- Strong hands-on expertise with Kubernetes (EKS preferred): cluster creation, deployments, scaling, and troubleshooting
- Solid understanding of networking fundamentals (VPCs, routing, DNS, load balancing, security groups)
- Experience with CI/CD pipelines, deployment tools, and infrastructure automation
- Working knowledge of databases and query optimization, and understanding how applications behave under load
- Familiarity with Kafka or other messaging systems
- Comfortable conducting code reviews and participating in coding focused interviews
- Strong operational mindset with experience in incident management and on-call rotations
- Clear communicator and collaborative teammate who values documentation and knowledge sharing
Technologies We Like
- AWS, Amazon EKS, Terraform, Jsonnet
- Kubernetes, Helm, CI/CD tooling
- Python (automation, tooling, reliability engineering)
- Prometheus, Grafana, logging and monitoring platforms
- PostgreSQL and other production databases
- Kafka or event driven systems
- Linux (Ubuntu or similar)
Benefits
We take care of you, so you can take care of business. We care about our people. That’s why we provide everything you—and your career—need to thrive at S&P Global.
- Health & Wellness: Health care coverage designed for the mind and body
- Flexible Downtime: Generous time off helps keep you energized
- Continuous Learning: Access a wealth of resources to grow your career
- Invest in Your Future: Competitive pay, retirement planning, continuing education with company-matched student loan contribution, and financial wellness programs
- Family Friendly Perks: Perks for partners and children, with strong benefits for families
- Beyond the Basics: Perks such as retail discounts and referral incentives
Equal Opportunity and Compliance
S&P Global is an equal opportunity employer and all qualified candidates will receive consideration without regard to race/ethnicity, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, marital status, veteran status, unemployment status, or any other status protected by law. Only electronic submissions will be considered. If you need an accommodation during the application process, please email EEO.Compliance@spglobal.com. US Candidates Only: Know Your Rights: Workplace discrimination is illegal.
For recruitment fraud awareness, see our guidelines and report suspicious activity to reportfraud@spglobal.com. Job ID: 326961. Location: Hyderabad, Telangana, India.