Manager - SRE

GreyOrange

Gurugram District

On-site

INR 4,500,000 - 6,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GreyOrange is seeking an experienced SRE Manager to lead a team responsible for reliability across mission-critical systems. You will own 24/7 on-call rotation, drive scalable infrastructure, and partner with engineering and product to embed reliability into the lifecycle.

The role requires strong leadership, cloud-native architecture expertise, and hands-on skills in Docker/Kubernetes, CI/CD, and observability. Excellent communication and stakeholder management are essential.

Qualifications

  • Willingness to work in a 24/7 operational environment.
  • 10–14 years of overall experience with 3–5 years in leadership/people management in SRE/DevOps.
  • Strong background in system design, distributed systems, cloud-native architecture.
  • Hands-on automation, containerization, CI/CD and config management with familiarity in observability.

Responsibilities

  • Lead and mentor a team of SREs, promoting accountability and technical excellence.
  • Design and own 24/7 on-call rotation and escalation matrix; lead critical production incident response.
  • Drive scalable, reliable, and secure infrastructure solutions with stakeholders.
  • Define and track SLIs/SLOs/SLAs; ensure adherence across services.
  • Oversee postmortems, root cause analysis, and continuous improvement initiatives.
  • Build automation frameworks to reduce toil and boost developer productivity.
  • Own capacity planning, disaster recovery, and performance optimization.

Skills

SRE leadership
Cloud architecture
Docker/Kubernetes
CI/CD automation
Observability
Incident management
SLI/SLO/SLAs
Networking/Linux
Go/Python/Bash
Leadership & stakeholder mgmt

Tools

Docker
Kubernetes
Ansible
Chef
CI/CD tools

Job description

About GreyOrange

GreyOrange is a global leader in AI-driven robotic automation software and hardware, transforming distribution and fulfillment centers worldwide. Our solutions increase productivity, empower growth and scale, mitigate labor challenges, reduce risk and time to market, and create better experiences for customers and employees. Founded in 2012, GreyOrange is headquartered in Atlanta, Georgia, with offices and partners across the Americas, Europe and Asia.

The SRE team at GreyOrange ensures the stability, availability, and scalability of mission‑critical production systems while driving incident management, automation, and operational excellence.

As an SRE Manager, you will lead a team of site reliability engineers, define best practices, and partner with engineering, product, and infrastructure teams to deliver resilient and scalable systems. You will be responsible for building a culture of reliability, ensuring operational readiness, and mentoring engineers to excel in reliability engineering.

Requirements
  • Willingness to work in a 247 operational environment
  • 10–14 years of overall experience, with at least 3–5 years in a leadership/people management role within SRE, DevOps, or Infrastructure Engineering.
  • Strong technical background in system design, distributed systems, and cloud‑native architecture.
  • Hands‑on expertise in automation, containerization (Docker, Kubernetes), CI/CD tools, and configuration management (Ansible, Chef, etc.).
  • Proficiency in programming/scripting languages (Python, Bash, Go, etc.) for automation and tooling.
  • Deep understanding of observability platforms (Grafana, Splunk, Dynatrace, Prometheus, etc.) and implementing monitoring strategies.
  • Proven experience with incident and problem management, including driving postmortems and continuous improvement.
  • Familiar with SLI, SLO, SLA, and Error Budget frameworks and applying them across multiple services.
  • Strong knowledge of networking, Unix/Linux systems, cloud platforms (AWS, GCP, or Azure), and databases.
  • Ability to drive reliability goals across multiple teams while balancing delivery speed and stability.
  • Excellent communication, leadership, and stakeholder management skills.
What You’ll Do
  • Lead and mentor a team of SREs, fostering a culture of accountability, collaboration, and technical excellence.
  • Design, own, and continuously improve the team’s 247 on‑call rotation and escalation matrix, ensuring sustainable coverage while personally stepping in as the point of escalation for critical production issues or as & when needed as a owner
  • Drive the design, implementation, and adoption of scalable, reliable, and secure infrastructure solutions.
  • Partner with engineering and product teams to embed reliability best practices into the development lifecycle.
  • Define and track reliability metrics (SLIs/SLOs/SLAs) and ensure adherence across services.
  • Oversee incident response, root cause analysis, and implement learnings to minimize recurrence.
  • Build automation frameworks to reduce operational toil and improve developer productivity.
  • Own capacity planning, disaster recovery strategies, and performance optimization of systems.
  • Act as the point of escalation for critical production issues and provide leadership during incidents.
  • Establish long‑term strategies for reliability, scalability, and cost efficiency of the platform.
  • Represent the SRE function with senior leadership, providing updates on reliability posture and improvement plans.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

GreyOrange • Gurugram District

On-site
INR 3,000,000 - 5,500,000
Site Reliability Engineer 2
Site Reliability Engineer 2

GreyOrange • Gurugram District

Hybrid
INR 2,600,000 - 4,600,000
Engineering Manager
Engineering Manager

WaferWire Cloud Technologies • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Lead SRE
Lead SRE

Cvent, Inc. • India

On-site
INR 2,500,000 - 4,500,000
SRE Program Manager
SRE Program Manager

Accenture in India • Maharashtra

On-site
INR 4,000,000 - 6,500,000
Lead SRE
Lead SRE

Baazi Games • New Delhi

On-site
INR 3,500,000 - 6,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Sr Engineering Manager
Sr Engineering Manager

Honeywell Technologies • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior SRE
Senior SRE

CloudRaft • India

On-site
INR 2,500,000 - 4,500,000
Competitive salary
Premium health insurance & wellness
AI stack & GPU infrastructure
+2