Engineering Manager

WaferWire Cloud Technologies

Hyderabad

On-site

INR 4,000,000 - 7,000,000

Full time

27 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

WaferWire Cloud Technologies seeks an SRE Manager to lead Reliability Engineering, Automation, and Operational Excellence. You will own the reliability architecture, monitor systems, and guide incident response at scale across multiple workstreams.

The role emphasizes proactive risk identification, AIOps adoption, and continuous improvement of cloud operations. The role requires hands-on leadership of a 15+ engineer team, strong collaboration with client leads, and a proven track record in

Qualifications

  • Hands-on SRE or platform engineering management leading 15+ engineers across geographies.
  • Proven reliability architecture ownership with code review rigor.
  • Track record of proactive improvements and delivery growth beyond scope.

Responsibilities

  • Own reliability architecture decisions for monitoring, alerting, automation, and IaC tooling.
  • Lead complex incident debugging and rapid root cause resolution.
  • Drive on-call health metrics, SLA adherence, and operational standards.

Skills

SRE
Platform engineering
Incident management
Leadership

Tools

KQL
Telemetry
Distributed tracing

Job description

SRE Manager – Reliability Engineering, Automation & Operational Excellence

Role Overview

Step into a technically hands-on SRE leadership role where you will engineer reliability – not just manage it. Own architecture for monitoring and automation, review infrastructure-as-code with rigor, author operational designs, and personally dive into complex incidents.

Bring an engineering-first mindset: be the leader who sees beyond what is currently working, identifies risks before they become incidents, champions AIOps, and actively expands the team into adjacent platform engineering.

You will lead a 15+ member team supporting large-scale cloud platform reliability across multiple workstreams – driving operational excellence, cloud cost efficiency, and engineering innovation. This is a high-visibility role as the single point of accountability between the vendor team and client engineering leads.

Key Responsibilities

Architecture & Technical Leadership

  • Own reliability architecture: drive decisions for monitoring, alerting, automation, and infrastructure tooling; conduct code reviews on IaC, scripts, and configs
  • Personally engage in complex incident debugging using KQL, telemetry, and distributed traces to unblock the team and drive rapid root cause resolution
  • Author runbook designs, define operational architecture, standardize procedures, and continuously raise the reliability bar

Delivery & Execution

  • Drive delivery in Agile/Scrum against contractual commitments: sprint progress, capacity management, delivery forecasting
  • Manage on-call rotation with health tracking (burnout, incident volume, team well-being)
  • Drive incident management excellence: SLA targets, post-incident reviews, MTTD/MTTR tracking, corrective actions, systemic improvements
  • Drive deployment operations: staged rollouts, Blue/Green deployments, deployment validation gates, safe deployment practices across cluster-based infrastructure

Cloud Cost & Operational Efficiency

  • Own cloud cost management: Azure spend monitoring, optimization, efficiency practices, burn rate reporting against budget
  • Drive AI and automation: AIOps adoption, automated incident detection, intelligent alerting, auto-remediation, Copilot-assisted engineering
  • Think beyond current operations: identify reliability risks before incidents, propose innovative automation, expand into adjacent platform engineering
  • Solution for new engagements: scoping, estimation, technical proposals, and delivery model shaping for new workstreams

People & Stakeholders

  • Drive multi-stakeholder coordination: primary technical interface with client leads, priority translation, platform roadmap awareness, proactive capacity positioning
  • Take full ownership: end-to-end accountability for team outcomes, proactive issue resolution, hold team to SLA and operational standards
  • Drive recruitment, professional development, and culture: technical interviews, team connects, mentor/coach across levels (intern to principal)
  • gDeep infrastructure debugging across distributed systems and cloud service degradatio
  • nSees beyond current operations to identify risks and propose innovative automatio
  • nSystems thinking for reliability architecture trade-off
  • sData-driven incident analysis for staffing, process, and tooling decision
  • sFinancial awareness for cloud spend management and team value demonstratio

n
Technic

  • alSRE/DevOps leadership, operational architecture, distributed systems (Service Fabric, Kubernetes/AK
  • S)Incident management, capacity planning, cloud cost optimizati
  • m)Monitoring/observability (KQL, telemetry, distributed tracing), AIOps, automation tooli
  • ngOn-call rotation design and health monitori
  • ngDelivery metrics (MTTD, MTTR, deployment frequency, SLA complianc

e)
Required Skills & Qualificati

  • onsHands-on SRE or platform engineering management leading 15+ engineers across geograph
  • iesProven reliability architecture ownership, code review rigor, and complex infrastructure debugg
  • ingTrack record of proactive improvements, solutioning for new engagements, and delivery growth beyond sc
  • opeExperience managing contract-based delivery with SLA accountability and operational commitme
  • ntsStrong design reviews, incident coordination, executive reporting, and stakeholder communicat
  • ionOwnership mindset delivering operational excellence beyond expectati

ons
Preferred Qualificat

  • ionsImproving MTTD/MTTR through technical and process cha
  • ngesCloud cost optimization and large-scale SaaS operat
  • ionsAIOps adoption (auto-remediation, intelligent alerting) in SRE t
  • eamsOn-call health management and sustainable support mo
  • delsSolutioning and estimation for new reliability engineering engagem
  • entsOperational architecture documents, runbook designs, reliability stand
  • ardsService Fabric, Kubernetes, or large-scale cluster-based platform operat

ions
Why You'll Love This

  • RoleEngineer reliability at scale – own architecture decisions that directly impact platform uptime and perfor
  • manceChampion AIOps, automation, and innovation while leading a high-performing team across multiple workst
  • reamsShape operational excellence as the trusted engineering partner to client leadership with direct business i
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Technologies Pvt. Ltd. • Pune District

On-site
INR 900,000 - 1,400,000
Lead Engineer - Reliability Engineering
Lead Engineer - Reliability Engineering

StoneX Group Inc. • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Visa Consolidated Support Services India • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Senior Software Engineer - Site Reliability
Senior Software Engineer - Site Reliability

Synthlane • Gurugram District

On-site
INR 3,000,000 - 5,700,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Augusta Infotech • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Maharashtra

On-site
INR 1,800,000 - 2,500,000
Lead SRE
Lead SRE

Cvent, Inc. • India

On-site
INR 2,500,000 - 4,500,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Namely • India

On-site
INR 1,500,000 - 2,500,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000