Sr Engineering Manager - SRE

OneAdvanced Limited

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

OneAdvanced Limited is seeking a Senior Manager of Site Reliability Engineering to own the monitoring and observability roadmap across Health, Legal, Education, and Workforce sectors. You will lead a 15–20 person SRE team, driving reliability, scalability, and AI-assisted incident resolution.

You will partner with Engineering, Security, Data, and Product teams to embed operational excellence from design to production and push for autonomous monitoring using AI agents.

Qualifications

  • 10+ years of experience in Site Reliability Engineering, observability, or platform engineering, including people leadership.
  • Proven track record leading teams of 15+ engineers.
  • Hands-on experience building and operating monitoring and observability platforms at scale.

Responsibilities

  • Own and drive the monitoring and observability roadmap across the product portfolio.
  • Lead, coach and inspire a high-performing team of SRE engineers.
  • Shape and deliver the SRE roadmap alongside the Head of Platform.

Skills

Leadership
Observability
Cloud infrastructure
AI agent deployment
Incident management
Mentoring
Cross-functional collaboration
Communication

Tools

Grafana
LogicMonitor
ServiceNow
Jira
AWS
Kubernetes
CI/CD

Job description

Job Summary

OneAdvanceds Site Reliability Engineering function is responsible for how well we can see into our own systems, and how fast we can act on what we see. As Senior Manager, Site Reliability Engineering, you’ll own the monitoring and observability roadmap across our full product portfolio - spanning Health, Legal, Education, Workforce Management, and other sectors - and lead a team of 15-20 SRE engineers to deliver it.

You’ll also push the function further: building and deploying AI agents that handle monitoring, alerting, and incident resolution directly - reducing how much of this work depends on a person watching a dashboard, and moving us toward a more autonomous operating model.

Mandate: To lead OneAdvanceds Site Reliability Engineering function - building the monitoring and observability roadmap that gives every product across Health, Legal, Education, Workforce Management, and other sectors real visibility into its own health, and increasingly uses AI agents to detect, alert on, and resolve incidents before they need a person.

Responsibilities
Monitoring Observability Roadmap
  • Own and drive the monitoring and observability roadmap across the product portfolio - Health, Legal, Education, Workforce Management, and other sectors.
  • Set the standard for what well-monitored means for a product, and hold the portfolio to it.
  • Evaluate, select, and evolve the monitoring toolset to match the scale and complexity of the estate.
  • Lead, coach and inspire a high-performing team of Site Reliability Engineers.
  • Shape and deliver our Site Reliability Engineering roadmap alongside the Head of Platform.
  • Champion modern engineering practices including SLIs, SLOs, error budgets, observability and automation.
  • Improve the reliability, scalability and performance of our cloud platforms and digital services.
  • Partner with Engineering, Security, Data and Product teams to embed operational excellence from design through to production.
  • Drive the adoption of our observability platform, helping teams gain deeper insight into the health and performance of their services.
  • Lead incident learning, continuous improvement and automation initiatives that reduce operational toil.
  • Provide technical leadership across AWS, Kubernetes, Infrastructure as Code, CI/CD and distributed system
AI-Driven Monitoring Incident Resolution
  • Design, build, and deploy AI agents that handle monitoring, alerting, and first-line incident resolution.
  • Identify where agentic automation can safely replace manual triage, and where human judgment still needs to stay in the loop.
  • Continuously improve agent accuracy and trust based on real incident outcomes.
Incident Management
  • Partner closely with Major Incident Management to ensure observability data drives faster detection and diagnosis during live incidents.
  • Drive root cause analysis for monitoring or alerting gaps that contributed to incident impact or delay.
  • Ensure post-incident learnings translate into concrete monitoring and alerting improvements.
Team Leadership
  • Lead, grow, and mentor a team of 10-15 SRE engineers, building deep observability and automation expertise.
  • Own hiring, performance management, and career development for the team.
  • Build a culture of ownership, curiosity, and continuous improvement.
Cross-Functional Partnership
  • Work closely with Engineering, Services, and Customer Success teams to ensure monitoring reflects what actually matters to product reliability and customer experience.
  • Represent SRE Observability in cross-functional planning and governance forums.
  • Act as an escape point for observability and monitoring gaps raised by any stakeholder team.
What You Will Have
  • 10+ years of experience in Site Reliability Engineering, observability, or platform engineering, including people leadership experience.
  • A proven track record leading teams of 15+ engineers.
  • Hands-on experience building and operating monitoring and observability platforms at scale.
  • Practical experience building or deploying AI agents for monitoring, alerting, or incident resolution - not just familiarity with the concept.
  • Experience partnering with incident management functions to improve detection and response.
  • Strong communication skills, comfortable working across engineering, services, and customer-facing teams.
  • Experience leading Site Reliability Engineering, Platform Engineering, DevOps or Cloud Infrastructure teams.
  • Strong technical expertise in cloud-native technologies, ideally within AWS, Azure and Private cloud platforms.
  • Experience operating and improving large-scale distributed systems.
  • A passion for coaching, mentoring and helping engineers thrive.
  • Experience driving observability, automation and operational excellence.
  • The ability to build strong relationships and influence stakeholders across Engineering, Product, Security and Data.
  • A pragmatic approach to balancing reliability, innovation and customer impact.

It would be great if you also had

  • Experience with monitoring and observability tools such as Grafana, LogicMonitor, or equivalent.
  • Experience with ServiceNow and Jira for incident and delivery tracking.
  • Hands-on experience with AI-assisted engineering tools (e.g. Claude Code, GitHub Copilot) beyond monitoring use cases.
  • Experience operating across multi-sector or multi-product portfolios.
  • Relevant certifications in observability platforms or cloud providers (AWS, Azure)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Engineering Manager - SRE
Senior Engineering Manager - SRE

OneAdvanced • Bengaluru

On-site
INR 3,500,000 - 7,500,000
Wellbeing programme
20 days annual leave
Employee Assistance Programme
+1
Lead Engineer - Reliability Engineering
Lead Engineer - Reliability Engineering

StoneX Group Inc. • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Namely • India

On-site
INR 1,500,000 - 2,500,000
SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Senior SRE Engineer
Senior SRE Engineer

EPAM Systems • Coimbatore District

Hybrid
INR 1,800,000 - 3,000,000
Senior SRE Engineer
Senior SRE Engineer

EPAM Systems • Maharashtra

Hybrid
INR 3,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

PwC Acceleration Center India • Bengaluru

On-site
INR 2,200,000 - 3,800,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Senior Consultant - Site Reliability Engineer
Senior Consultant - Site Reliability Engineer

Darwinbox Digital Solutions Pvt. Ltd. • Hyderabad

On-site
INR 3,000,000 - 5,200,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Bengaluru

On-site
INR 2,500,000 - 4,200,000