Senior Platform SRE

IG Infotech

Bengaluru

On-site

INR 1,200,000 - 1,600,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

IG Infotech in Bengaluru is seeking a Platform SRE to enhance reliability across engineering teams. The role involves building platforms, implementing monitoring solutions, and mentoring junior SREs. Candidates should have experience with OpenTelemetry, Kubernetes, and production systems.

This position offers opportunities to improve system reliability and engage in high-throughput environments. Join a team dedicated to operational excellence and reliability practices!

Qualifications

  • Hands-on experience with OpenTelemetry and ability to instrument Java or Python services.
  • Track record designing customer-meaningful SLIs and setting error budgets.
  • Knowledge of container orchestration tools like Kubernetes required.

Responsibilities

  • Build and own the reliability platform for engineers.
  • Implement monitoring and observability using OpenTelemetry.
  • Maintain SLOs, error budgets, and operational readiness.

Skills

OpenTelemetry
Datadog
Kubernetes
Python
Java
Terraform

Tools

Honeycomb
Dynatrace
Grafana

Job description

Role Overview

The Platform SRE team is the engine of IG’s reliability programme. Working within Infrastructure & Operations, we build the platform, standards, and tooling that make reliability the default for every engineering team at IG.

Responsibilities
  • Build and own the reliability platform that hundreds of engineers depend on.
  • Implement comprehensive monitoring and observability using OpenTelemetry and distributed tracing.
  • Maintain SLOs, error budgets and burn‑rate tracking.
  • Establish and maintain 24/7 operational readiness, including automated deployments, blue/green releases, and zero‑downtime patching strategies.
  • Engineer self‑healing capabilities such as auto‑remediation, error‑budget‑gated rollback, and automated traffic rerouting.
  • Design and run chaos experiments across the AWS estate to turn failure scenarios into engineering improvements.
  • Build automation tools and CI/CD pipelines that embed reliability practices, while applying software engineering discipline including version control, code reviews, and testing.
  • Contribute to the SRE AI agent, IG’s agentic tooling for incident investigation and reliability review, built on AWS frontier models.
  • Mentor junior SREs and Reliability Champions on reliability patterns and production engineering discipline.
  • Set and uphold standards, author and evolve the SRE standards that underpin the Guild: SLO methodology, error budget policy, observability instrumentation guide, and Production Readiness Review (PRR) checklist.
  • Mentor developers on reliability patterns such as circuit breakers, retry logic, and fault tolerance.
  • Work with development teams and Reliability Champions to design SLOs on customer journeys rather than per‑service.
  • Assist and guide teams in system design, capacity planning, architectural reviews and closing observability gaps.
  • Own incident response and learning: facilitate blameless post‑incident reviews (PIRs) within five working days, maintain the Lessons Register, track remediation actions to closure, and surface patterns across incidents quarterly.
  • Participate in the SRE Guild, mentor engineers across the organisation, and help define what good looks like in production.
Qualifications
  • Hands‑on experience with OpenTelemetry, Honeycomb, Datadog, Dynatrace, or Grafana, and the ability to instrument Java or Python services.
  • Proven track record designing customer‑meaningful SLIs, setting error budgets, configuring multi‑window burn‑rate alerts, and working with development teams on reliability measurement.
  • Experience building pipelines with safety mechanisms: blue/green and canary releases, automated rollback, and DORA metrics integration.
  • Knowledge of container orchestration: Kubernetes (EKS, AKS, or GKE) required; HashiCorp Nomad is a strong advantage.
  • Solid understanding of cloud networking and IaC (Terraform preferred).
  • Production‑quality coding in Java and/or Python; comfortable contributing to application codebases to implement reliability patterns.
  • Strong understanding of distributed systems, including circuit breakers, bulkheads, idempotency, graceful degradation, and load‑shedding; experience in high‑throughput, low‑latency environments is preferred.
  • On‑call experience on production systems, blameless PIR facilitation, contributing‑factor analysis, and driving action items to closure; familiarity with PagerDuty and ServiceNow helpful.
  • Experience designing and executing hypothesis‑driven chaos experiments with blast‑radius controls and gap‑to‑impact tolerance analysis; experience with AWS FIS, Gremlin, or equivalent.
  • Comfortable writing RFCs, presenting at engineering forums, and building standards that others will adopt.
Experience Requirements
  • Track record in high‑throughput, production environments (financial services, trading platforms, or similar mission‑critical systems preferred).
  • Demonstrated ability to improve system reliability and performance at scale.
  • Experience working collaboratively with development teams to implement observability and reliability improvements.
  • Strong troubleshooting skills in distributed systems environments.
Core Competencies
  • Systems thinking approach to problem‑solving.
  • Excellent communication skills for cross‑functional collaboration and technical enablement.
  • Ability to balance hands‑on development work with operational responsibilities.
  • Strong bias toward automation and eliminating manual toil.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Cybage Software • Pune District

On-site
INR 2,500,000 - 4,000,000
Lead Engineer - Reliability Engineering
Lead Engineer - Reliability Engineering

StoneX Group Inc. • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
Site Reliability Engineer
Site Reliability Engineer

Epam Systems • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Engineering Manager
Engineering Manager

WaferWire Cloud Technologies • Hyderabad

On-site
INR 4,000,000 - 7,000,000
SRE Expert
SRE Expert

HCLTech • Bengaluru

On-site
INR 1,500,000 - 2,400,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
DevOps & SRE Lead
DevOps & SRE Lead

Syngentagroup • Pune District

On-site
INR 4,000,000 - 7,000,000
Senior Manager System Reliability Engineering
Senior Manager System Reliability Engineering

GE Vernova, Inc. • Hyderabad

On-site
INR 4,000,000 - 6,000,000
Relocation Assistance Provided: Yes