Director – Site Reliability Engineering (SRE)

Umanist NA

Navi Mumbai

On-site

INR 6,000,000 - 12,000,000

Full time

17 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Umanist NA is seeking a Director of Site Reliability Engineering to lead reliability and operational excellence across multiple SaaS products. The candidate should combine strong software engineering fundamentals with deep SRE expertise, cloud infrastructure, observability, and CI/CD.

The leader will build scalable reliability programs, improve availability, establish SLI/SLO practices, and drive operational excellence through measurable metrics and KPIs while mentoring senior engineers and

Qualifications

  • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field.
  • 18+ years of experience in Software Engineering, SRE, Site Reliability, Platform Engineering, or related roles.
  • 5+ years of leadership experience at Director level or equivalent.
  • Strong experience in SaaS / B2B SaaS / cloud product companies.
  • Proven ability to apply software engineering principles to solve reliability and operational challenges.

Responsibilities

  • Lead and develop SRE / Reliability Engineering teams supporting multiple SaaS products.
  • Establish and drive reliability engineering strategy across the organization.
  • Define and manage SLIs, SLOs, SLAs, error budgets, and reliability KPIs.
  • Drive observability, monitoring, alerting, and proactive performance management.
  • Establish and improve incident response, escalation, troubleshooting, and post-incident review processes.
  • Partner with Software Engineering, Product, Cloud/Platform, Security, and Infrastructure teams.
  • Apply software engineering principles to automate and solve reliability and operational challenges.
  • Lead reliability programs across multiple SaaS products.

Skills

SRE Leadership
Kubernetes
AWS
CI/CD
Observability
Distributed Systems
Incident Management

Education

Bachelor's degree in Computer Science, Information Systems, Engineering, or related field

Tools

Terraform
Prometheus
Grafana
Datadog
New Relic
Splunk
Docker
Kubernetes

Job description

Job Title: Director – Site Reliability Engineering (SRE)

Industry: Candidate should be from B2B SaaS / Cloud Product companies in recent work experience

Experience: 18+ Years

Leadership: 5+ Years at Director / Senior Engineering Leadership Level(super-mandate, non-negotiable)

Notice Period: Immediate to 30 Days

Job Location: Hyderabad

Role Overview

We are looking for an experienced Director of Site Reliability Engineering (SRE) to lead reliability and operational excellence across multiple SaaS products.

The ideal candidate combines strong software engineering fundamentals with deep expertise in SRE, cloud infrastructure, observability, monitoring, incident management, CI/CD, and distributed systems.

This leader will be responsible for building scalable reliability programs, improving availability and performance, establishing SLI/SLO practices, and driving operational excellence through measurable metrics and KPIs.

Key Responsibilities
  • Lead and develop SRE / Reliability Engineering teams supporting multiple SaaS products.
  • Establish and drive reliability engineering strategy across the organization.
  • Define and manage SLIs, SLOs, SLAs, error budgets, and reliability KPIs.
  • Drive observability, monitoring, alerting, and proactive performance management.
  • Establish and improve incident response, escalation, troubleshooting, and post-incident review processes.
  • Partner with Software Engineering, Product, Cloud/Platform, Security, and Infrastructure teams.
  • Apply software engineering principles to automate and solve reliability and operational challenges.
  • Improve system availability, scalability, resilience, performance, and operational efficiency.
  • Drive CI/CD improvements and deployment reliability.
  • Architect and support applications and infrastructure running on high-growth cloud platforms.
  • Lead reliability programs across multiple B2B SaaS products.
  • Establish engineering metrics and use data/KPIs to identify and resolve operational gaps.
  • Drive automation and reduction of repetitive operational work.
  • Mentor engineering leaders and senior SRE engineers.
  • Communicate reliability strategy, risks, metrics, and business impact to senior leadership.
  • Influence engineering teams and cross-functional stakeholders to adopt reliability best practices.

Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field.

  • 18+ years of experience in Software Engineering, SRE, Site Reliability, Platform Engineering, or related reliability/engineering roles.
  • 5+ years of leadership experience at Director level or equivalent.
  • Strong experience in SaaS / B2B SaaS / cloud product companies.
  • Proven ability to apply software engineering principles and practices to solve reliability and operational challenges.
  • Strong expertise in SLI/SLO, monitoring, observability, and reliability engineering.
  • Strong experience with CI/CD and modern software delivery practices.
  • Strong experience with incident response, problem management, RCA, and production operations.
  • Strong AWS expertise.
  • Experience with container orchestration, such as Kubernetes.
  • Experience leading reliability programs across multiple SaaS products.
  • Experience architecting applications or infrastructure for high-growth cloud platforms.
  • Experience with large-scale distributed systems in B2B SaaS environments.
  • Strong leadership, communication, stakeholder management, and influencing skills.
  • Demonstrated experience driving operational excellence through metrics, KPIs, SLOs, and reliability objectives.
Preferred Skills
  • Kubernetes / containerized environments
  • AWS cloud architecture
  • Infrastructure and application observability
  • Distributed systems
  • Microservices
  • Infrastructure automation
  • Infrastructure-as-Code
  • Terraform
  • Prometheus / Grafana
  • Datadog / New Relic / Splunk or similar observability platforms
  • CI/CD platforms
  • Disaster recovery and business continuity
  • Capacity and performance engineering
  • Chaos/resilience engineering

Skills: cloud,saas,b2b,monitoring,sli/slo,production operations,kubernetes,ci/cd,container orchestration,leadership,metrics,multiple saas products,aws,large-scale distributed systems,problem management,observability,reliability engineering,kpis,b2b saas environments,slos

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director – Site Reliability Engineering (SRE)
Director – Site Reliability Engineering (SRE)

Umanist NA • Mumbai

On-site
INR 6,000,000 - 9,000,000
Director – Site Reliability Engineering (SRE)
Director – Site Reliability Engineering (SRE)

Umanist NA • Hyderabad

On-site
INR 3,500,000 - 5,200,000
Senior Consultant - Site Reliability Engineer
Senior Consultant - Site Reliability Engineer

HCA Healthcare • Hyderabad

On-site
INR 3,000,000 - 5,200,000
Software Engineer
Software Engineer

PwC • Hyderabad, Bengaluru

Hybrid
INR 2,800,000 - 5,200,000
Site Reliability Engineer
Site Reliability Engineer

Mumba Technologies, Inc. • Gurugram District

Hybrid
INR 1,500,000 - 2,100,000
Site Reliability Engineering (SRE) Manager
Site Reliability Engineering (SRE) Manager

Acesoft Labs • Hyderabad

Hybrid
INR 4,000,000 - 6,000,000
Tech & Digital-Lead Site Reliability Engineer
Tech & Digital-Lead Site Reliability Engineer

Hdfc Bank • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Acesoft Labs • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

Acesoft Labs • Ahmedabad District

Hybrid
INR 400,000 - 700,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VMC Soft Technologies, Inc • Hyderabad

On-site
INR 1,500,000 - 2,000,000