Director – Site Reliability Engineering (SRE)

Umanist NA

Hyderabad

On-site

INR 4,000,000 - 7,000,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Umanist NA in Hyderabad, India, seeks a Director of Site Reliability Engineering to lead reliability across multiple B2B SaaS products. The role requires deep SRE expertise, cloud infrastructure, observability, incident management, CI/CD, and distributed systems.

You will build scalable reliability programs, set SLIs/SLOs, improve availability and performance, mentor senior engineers, and report metrics to leadership.

Qualifications

  • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field.
  • 18+ years of experience in Software Engineering, SRE, Site Reliability, Platform Engineering, or related reliability/engineering roles.
  • 5+ years of leadership experience at Director level or equivalent.
  • Strong experience in SaaS / B2B SaaS / cloud product companies.
  • Proven ability to apply software engineering principles and practices to solve reliability and operational challenges.
  • Strong expertise in SLI/SLO, monitoring, observability, and reliability engineering.
  • Strong experience with CI/CD and modern software delivery practices.
  • Strong experience with incident response, problem management, RCA, and production operations.
  • Strong AWS expertise.
  • Experience with container orchestration, such as Kubernetes.
  • Experience leading reliability programs across multiple SaaS products.
  • Experience architecting applications or infrastructure for high-growth cloud platforms.
  • Experience with large-scale distributed systems in B2B SaaS environments.
  • Strong leadership, communication, stakeholder management, and influencing skills.
  • Demonstrated experience driving operational excellence through metrics, KPIs, SLOs, and reliability objectives.

Responsibilities

  • Lead and develop SRE/ Reliability Engineering teams supporting multiple SaaS products.
  • Establish and drive reliability engineering strategy across the organization.
  • Define and manage SLIs, SLOs, SLAs, error budgets, and reliability KPIs.
  • Drive observability, monitoring, alerting, and proactive performance management.
  • Establish and improve incident response, escalation, troubleshooting, and post-incident review processes.
  • Partner with Software Engineering, Product, Cloud/Platform, Security, and Infrastructure teams.
  • Apply software engineering principles to automate and solve reliability and operational challenges.
  • Improve system availability, scalability, resilience, performance, and operational efficiency.
  • Drive CI/CD improvements and deployment reliability.
  • Architect and support applications and infrastructure running on high-growth cloud platforms.
  • Lead reliability programs across multiple SaaS products.
  • Establish engineering metrics and use data/KPIs to identify and resolve operational gaps.
  • Drive automation and reduction of repetitive operational work.
  • Mentor engineering leaders and senior SRE engineers.
  • Communicate reliability strategy, risks, metrics, and business impact to senior leadership.
  • Influence engineering teams and cross-functional stakeholders to adopt reliability best practices.

Skills

SRE leadership
Cloud architecture
Observability
CI/CD automation
Distributed systems
Kubernetes
Strong communication
Stakeholder management

Education

Bachelor's degree in Computer Science or related field

Tools

AWS
Kubernetes
Terraform
Prometheus/Grafana
Datadog/New Relic/Splunk
CI/CD platforms

Job description

Job Title: Director – Site Reliability Engineering (SRE)

Industry: B2B SaaS / Cloud Product

Experience: 18+ Years

Leadership: 5+ Years at Director / Senior Engineering Leadership Level

Notice Period: Immediate to 30 Days

Role Overview

We are looking for an experienced Director of Site Reliability Engineering (SRE) to lead reliability and operational excellence across multiple SaaS products.

The ideal candidate combines strong software engineering fundamentals with deep expertise in SRE, cloud infrastructure, observability, monitoring, incident management, CI/CD, and distributed systems.

This leader will be responsible for building scalable reliability programs, improving availability and performance, establishing SLI/SLO practices, and driving operational excellence through measurable metrics and KPIs.

Key Responsibilities
  • Lead and develop SRE / Reliability Engineering teams supporting multiple SaaS products.
  • Establish and drive reliability engineering strategy across the organization.
  • Define and manage SLIs, SLOs, SLAs, error budgets, and reliability KPIs.
  • Drive observability, monitoring, alerting, and proactive performance management.
  • Establish and improve incident response, escalation, troubleshooting, and post-incident review processes.
  • Partner with Software Engineering, Product, Cloud/Platform, Security, and Infrastructure teams.
  • Apply software engineering principles to automate and solve reliability and operational challenges.
  • Improve system availability, scalability, resilience, performance, and operational efficiency.
  • Drive CI/CD improvements and deployment reliability.
  • Architect and support applications and infrastructure running on high-growth cloud platforms.
  • Lead reliability programs across multiple B2B SaaS products.
  • Establish engineering metrics and use data/KPIs to identify and resolve operational gaps.
  • Drive automation and reduction of repetitive operational work.
  • Mentor engineering leaders and senior SRE engineers.
  • Communicate reliability strategy, risks, metrics, and business impact to senior leadership.
  • Influence engineering teams and cross-functional stakeholders to adopt reliability best practices.
Mandatory Requirements
  • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field.
  • 18+ years of experience in Software Engineering, SRE, Site Reliability, Platform Engineering, or related reliability/engineering roles.
  • 5+ years of leadership experience at Director level or equivalent.
  • Strong experience in SaaS / B2B SaaS / cloud product companies.
  • Proven ability to apply software engineering principles and practices to solve reliability and operational challenges.
  • Strong expertise in SLI/SLO, monitoring, observability, and reliability engineering.
  • Strong experience with CI/CD and modern software delivery practices.
  • Strong experience with incident response, problem management, RCA, and production operations.
  • Strong AWS expertise.
  • Experience with container orchestration, such as Kubernetes.
  • Experience leading reliability programs across multiple SaaS products.
  • Experience architecting applications or infrastructure for high-growth cloud platforms.
  • Experience with large-scale distributed systems in B2B SaaS environments.
  • Strong leadership, communication, stakeholder management, and influencing skills.
  • Demonstrated experience driving operational excellence through metrics, KPIs, SLOs, and reliability objectives.
Preferred Skills
  • Kubernetes / containerized environments
  • AWS cloud architecture
  • Infrastructure and application observability
  • Distributed systems
  • Microservices
  • Infrastructure automation
  • Infrastructure-as-Code
  • Terraform
  • Prometheus / Grafana
  • Datadog / New Relic / Splunk or similar observability platforms
  • CI/CD platforms
  • Disaster recovery and business continuity
  • Capacity and performance engineering
  • Chaos/resilience engineering

Skills: cd,cloud,ci,leadership,b2b,infrastructure,reliability,saas,software,reliability engineering

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Principal SRE
Principal SRE

Lloyds Banking Group • Hyderabad

Hybrid
INR 3,000,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Lead SRE
Lead SRE

Cvent, Inc. • India

On-site
INR 2,500,000 - 4,500,000
Site Reliability Engineer Lead (Immediate Joiner)
Site Reliability Engineer Lead (Immediate Joiner)

HiLabs • Pune District

On-site
INR 2,500,000 - 4,200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Lead SRE
Lead SRE

Cvent • Gurugram District

On-site
INR 4,000,000 - 7,000,000
SRE Program Manager
SRE Program Manager

Accenture in India • Maharashtra

On-site
INR 4,000,000 - 6,500,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Chennai District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000