Azure Cloud Lead - SRE

ViaPlus

Hyderabad

On-site

INR 2,000,000 - 3,000,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

ViaPlus is looking for a Lead Cloud SRE to ensure the reliability and performance of mission-critical platforms on Microsoft Azure. This hands-on role entails managing complex, distributed systems and leading incident responses while optimizing observability and reliability engineering.

Ideal candidates should have over 12 years of relevant experience, proficiency with Azure, and the ability to mentor teams. This position offers a dynamic work environment focused on innovation and excellence.

Qualifications

  • 12+ years of experience in production support / SRE roles.
  • Strong hands-on experience with Microsoft Azure.
  • Ability to mentor L1/L2 teams.

Responsibilities

  • Diagnose and resolve complex Azure VM issues.
  • Support and troubleshoot microservices-based architectures.
  • Lead P0/P1 incident bridges and drive resolution.

Skills

Production support experience
Microsoft Azure
Linux internals troubleshooting
Network fundamentals
Incident management
Mentoring teams

Education

B.E / B.Tech, MCA or equivalent degree
Microsoft Certified: DevOps Engineer Expert (AZ-400)
Microsoft Certified: Solutions Architect Expert (AZ-305)

Job description

Lead Cloud SRE

ViaPlus is seeking a Lead Cloud SRE to own the reliability, availability, and performance of large-scale, mission‑critical platforms running on Microsoft Azure. This role is responsible for maintaining production stability across complex, distributed systems by leading incident response, observability, and reliability engineering initiatives.

The Lead Cloud SRE will work hands‑on with Azure infrastructure, Kubernetes‑based and VM‑hosted microservices, networking, and data platforms to diagnose and resolve high‑severity production issues. The role involves deep root‑cause analysis using telemetry from Azure Monitor, Application Insights, and Log Analytics, as well as driving long‑term remediation through automation, architectural improvements, and systemic fixes.

Job Responsibilities
  • Diagnose and resolve complex Azure VM issues including boot failures, performance degradation, disk I/O latency, and memory leaks.
  • Troubleshoot VM Scale Sets, OS‑level issues across Linux and Windows, and patching or upgrade failures.
  • Analyze and remediate network connectivity issues involving NSGs, UDRs, DNS resolution, and routing configurations.
2. Application & Microservices Reliability
  • Support and troubleshoot microservices‑based architectures hosted on AKS and virtual machines.
  • Identify and resolve inter‑service latency, timeouts, retry storms, and cascading failure scenarios.
  • Diagnose application‑level issues such as thread pool exhaustion, memory leaks, misconfigurations, and resource contention.
  • Eliminate certificate, authentication, and upstream/downstream dependency failures impacting service availability.
  • Maintain and restore Service Fabric cluster health and stability.
  • Troubleshoot node failures, replica movement delays, quorum loss, and partition health issues.
  • Investigate upgrade and rollback failures, ensuring minimal service disruption.
  • Analyze and optimize both stateful and stateless service behaviours.
4. Traffic Management, Load Balancing & Edge Services
  • Troubleshoot HTTP 502/503/504 errors and backend pool health issues.
  • Debug probe failures, SSL/TLS termination, listener configurations, and routing rules.
  • Optimize WAF rules for security, performance, and reduced false positives.
  • Diagnose routing, caching, latency issues, and WAF‑related traffic blocks.
  • Investigate backend connectivity, health probes, and geo‑routing behaviour.
  • NGINX / Reverse Proxies: Debug connection resets, upstream timeouts, worker exhaustion.
  • Tune timeouts, buffers, keep‑alive settings, and load‑balancing strategies for high availability.
5. Database & Data Layer Reliability
  • Troubleshoot Azure SQL, Managed Instances, PostgreSQL, MySQL, and Cosmos DB.
  • Analyze slow queries, deadlocks, connection pool exhaustion, and resource contention.
  • Manage failovers, replication lag, throttling issues (DTU/RU limits), and high availability scenarios.
  • Collaborate on query optimization, execution plans, and indexing strategies.
6. Observability, Monitoring & Incident Management
  • Perform deep‑dive analysis using Azure Monitor, Application Insights, and Log Analytics (KQL).
  • Correlate infrastructure, application, and dependency telemetry to identify root causes.
  • Lead P0/P1 incident bridges, driving coordinated resolution under pressure.
  • Produce blameless root cause analyses (RCAs) with actionable corrective and preventive measures.
7. Site Reliability Engineering (SRE) Practices
  • Define, measure, and continuously improve SLIs, SLOs, and error budgets.
  • Drive automation and self‑healing solutions to improve service reliability.
  • Improve deployment safety through blue‑green, canary, and progressive delivery strategies.
  • Continuously reduce MTTR and eliminate recurring incidents through systemic fixes.
Skill Set
  • 12+ years of experience in production support / SRE roles
  • Strong hands‑on experience with Microsoft Azure
  • Expert‑level troubleshooting of Linux internals & networking fundamentals
  • Experience handling high‑severity production incidents
  • Security exposure (Azure WAF, Defender for Cloud)
  • Ability to mentor L1/L2 teams
Qualifications
  • Any Graduate with B.E / B.Tech, MCA or equivalent degree with more than 12+ years relevant work experience.
  • Microsoft Certified: DevOps Engineer Expert (AZ‑400)
  • Microsoft Certified: Solutions Architect Expert (AZ‑305)

ViaPlus is an equal opportunity employer and is committed to building a diverse and inclusive workplace. All qualified applicants will receive consideration for employment without regard to race, ethnicity, religion, gender, sexual orientation, national origin, age, disability, marital status, or any other characteristic protected by applicable law.

We invite you to explore this opportunity to be a part of the ViaPlus family.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Application Support Specialist - SRE
Application Support Specialist - SRE

ViaPlus • Hyderabad

On-site
INR 900,000 - 1,300,000
Assistant Manager - Azure Site Reliability Engineer
Assistant Manager - Azure Site Reliability Engineer

Promaynov Advisory Services Pvt. Ltd • Bengaluru

On-site
INR 1,400,000 - 2,100,000
Azure Site Reliability Engineer (SRE) + SQL - SaaS Operations
Azure Site Reliability Engineer (SRE) + SQL - SaaS Operations

Zensar Technologies • Pune District

On-site
INR 2,000,000 - 2,800,000
Site Reliability Engineer
Site Reliability Engineer

Finthrive • Gurugram District

On-site
INR 1,800,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

PwC India • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineer (SRE) L2
Site Reliability Engineer (SRE) L2

DigitalXNode • Ahmedabad District

On-site
INR 1,200,000 - 1,800,000
SRE - AWS, GCP & Azure
SRE - AWS, GCP & Azure

PibyThree • Navi Mumbai

On-site
INR 1,200,000 - 1,800,000
SRE - AWS, GCP & Azure
SRE - AWS, GCP & Azure

PibyThree • Thane

On-site
INR 1,200,000 - 1,500,000
Azure Network Specialist
Azure Network Specialist

ViaPlus • Hyderabad

On-site
INR 2,400,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Embarkgcc Services • Bengaluru

On-site
INR 1,200,000 - 1,800,000