Senior Site Reliability Engineering Manager

Seven N Half

Bengaluru

On-site

INR 2,000,000 - 2,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading tech company in Bengaluru seeks a Senior Site Reliability Engineering Manager to oversee cloud stability and automate infrastructure management. The ideal candidate has over 10 years of experience in cloud operations, especially in Azure environments, with a focus on reliability and performance optimization. This role involves leading teams towards automation and high availability standards.

Qualifications

  • Experience with Azure environments focusing on high reliability.
  • Hands-on expertise in cloud automation tools and practices.
  • Understanding of incident management and recovery processes.

Responsibilities

  • Lead the design and implementation of strategies for high availability.
  • Oversee the Cloud Infra Lifecycle Management Team.
  • Manage and mentor the SRE team for continuous improvement.

Skills

Cloud automation
SRE principles
Leadership
Problem-solving
Azure services management

Education

10+ years in cloud operations or SRE

Tools

Terraform
PowerShell
Azure DevOps
Log Analytics

Job description

Senior Site Reliability Engineering Manager
Senior Site Reliability Engineering Manager

Get AI-powered advice on this job and more exclusive features.

Direct message the job poster from Seven N Half

About Us:

The Super-app is a future-ready company that focuses on creating consumer-centric, high-engagement digital products. By creating a holistic presence across various touchpoints, we aim to be the trusted partner of every consumer and delight them by powering a rewarding life. The company's debut offering is a super-app that provides an integrated rewards experience across various consumer categories like groceries, fashion and electronics, travel and hospitality, health and fitness, entertainment, and financial services on a single platform.

Our Culture:

We cultivate a culture of innovation, inclusion for all employees and respect their individual strengths, views, and experiences. We thrive on the diversity of our talent in all forms and see it as a strength in building high performance teams across brands. As we rewrite commerce in India, change is the only constant in our day to day lives.

Role Overview:

We are looking for a Sr. Engineering Manager - SRE to oversee the stability, scalability, and delivery of our production environment, leveraging software engineering principles and automation to improve cloud infrastructure management and reduce operational costs. This role will play a key part in transitioning from manual processes to automated solutions by leading our current DevOps teams:

  • Cloud Infra Lifecycle Management Team: Focused on automated provisioning, capacity planning, and maintenance across all cloud platforms for production applications.
  • Cloud Infra Support Team: Responsible for supporting internal users with production and development environment requests, with a long-term goal of eliminating manual intervention through automation.

This role is ideal for a leader with a deep understanding of Azure cloud environments, SRE best practices, and a strong background in building automation-first operational models.

Key Responsibilities:

Stability, Scalability & Availability:

  • Lead the design and implementation of strategies to ensure high availability, reliability, and performance of production systems.
  • Apply lifecycle management techniques, including monitoring, capacity planning, and automated scaling, to cloud environments.
  • Establish Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical applications.

Cloud Lifecycle Management:

  • Oversee the Cloud Infra Lifecycle Management Team to build scalable, automated cloud provisioning workflows and optimize capacity.
  • Implement infrastructure-as-code (IaC) practices using tools like Terraform, PowerShell, and Azure Resource Manager (ARM) templates.
  • Ensure efficient cloud resource utilization and cost management strategies.

Cloud Support Operations:

  • Manage the Cloud Infra Support Team responsible for handling internal user requests related to production and development environments.
  • Develop efficient workflows for incident response and request resolution, with automation as the default approach.
  • Work towards eliminating the need for manual support teams by creating self-service solutions for internal users.

Automation & Transformation:

  • Lead the transition of manual processes to cloud automation through training, upskilling, and process reengineering.
  • Champion the use of automation to handle repetitive operational tasks, including monitoring, remediation, and deployments.
  • Foster a "first principles thinking" culture focused on engineering excellence and process simplification.
  • Build robust monitoring systems using Azure Monitor, Log Analytics, and Application Insights for proactive performance management.
  • Oversee incident response processes, ensuring rapid recovery and root cause analysis for production disruptions.
  • Implement disaster recovery and high-availability strategies across environments.

Security & Compliance:

  • Ensure all environments follow cloud security best practices, regulatory compliance, and corporate governance policies.
  • Manage identity and access controls, network security, and risk mitigation strategies.
  • Drive ongoing improvements in system resilience, operational efficiency, and service quality through automation and best practices.
  • Conduct regular performance reviews and capacity planning exercises to maintain optimal system health.
  • Provide coaching and mentorship to the SRE team, fostering a culture of continuous learning and technical excellence.
  • Lead efforts to upskill the team in cloud scripting, automation development, and site reliability best practices.

Reporting & Metrics:

  • Maintain detailed operational documentation and generate regular reports on system performance, reliability improvements, and cost efficiency efforts.

Basic Qualifications:

  • 10+ years of experience in cloud operations or SRE, with a strong focus on Azure environments.
  • Extensive experience in managing and optimizing Azure services like Virtual Machines, App Services, SQL Database, Networking, and Storage.
  • Hands-on expertise with cloud automation and IaC tools (Terraform, PowerShell, ARM templates, or Azure Automation).
  • Strong understanding of SRE principles, including error budgets, SLOs, SLIs, and incident management practices.
  • Proficiency with Azure DevOps and CI/CD pipeline management.
  • Expertise in cloud cost management and optimization.
  • Familiarity with monitoring, logging, and observability tools (e.g., Azure Monitor, Log Analytics, Security Centre).
  • Knowledge of Azure security practices, including identity and access management, firewalls, and compliance requirements.

Preferred Qualifications:

  • Experience managing hybrid or multi-cloud environments.
  • Experience implementing self-service workflows and internal user support automation.

Soft Skills:

  • Strong leadership and team management abilities.
  • Excellent communication and client engagement skills.
  • Analytical mindset with a proactive approach to problem-solving.
  • Ability to handle high-pressure situations with professionalism.
Seniority level
  • Seniority level
    Director
Employment type
  • Employment type
    Full-time
Job function
  • Job function
    Information Technology and Engineering
  • Industries
    IT Services and IT Consulting, Internet Marketplace Platforms, and Software Development

Referrals increase your chances of interviewing at Seven N Half by 2x

Sign in to set job alerts for “Site Reliability Engineering Manager” roles.
Engineering Manager - Core Search Engineering
Engineering Manager - Core Banking (SDE 5)
Engineering Manager, Power and Performance
Engineering Manager, Real World Journeys, Search

We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Engineering Manager, SRE [T500-18109]
Senior Engineering Manager, SRE [T500-18109]

ANSR • Hyderabad

On-site
INR 2,189,000 - 3,503,000
Sr. Lead Site Reliability Engineer – Technical & People Leadership
Sr. Lead Site Reliability Engineer – Technical & People Leadership

Shell Recharge Solutions • Bengaluru

On-site
INR 2,000,000 - 2,500,000
Flexible scheduling
Generous holiday package
Medical benefits
Senior Engineering Manager, SRE [T500-17089]
Senior Engineering Manager, SRE [T500-17089]

ANSR • Hyderabad

On-site
INR 7,000,000 - 8,000,000
Site Reliability Engineer
Site Reliability Engineer

Bounteous • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Lead Site Reliability Engineer [T500-19357]
Lead Site Reliability Engineer [T500-19357]

Deutsche Börse • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Azure Infrastructure & Operations Engineer
Azure Infrastructure & Operations Engineer

Idyllic Services • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Engineering Manager
Engineering Manager

Worko • Bengaluru

On-site
INR 8,500,000 - 10,000,000
Associate Engineer II - Site Reliability Engineer [T500-17376]
Associate Engineer II - Site Reliability Engineer [T500-17376]

ANSR • Hyderabad

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer
Site Reliability Engineer

LSEG • Bengaluru

On-site
INR 600,000 - 900,000
Competitive salary and benefits
Opportunities for learning and career development
Paid volunteering days
+1
Interesting Job Opportunity: Sciactive Solutions - Senior Software Engineering Manager - Full S[...]
Interesting Job Opportunity: Sciactive Solutions - Senior Software Engineering Manager - Full S[...]

Sciative - We Price Right • Navi Mumbai

On-site
INR 2,500,000 - 3,500,000