SRE Manager
Date: 4 Sept 2026 Company: QualityAI Country/Region: IN
SRE Lead, Platform Reliability and OperationsNewLeaf Azure application platform
This is the role specification for a senior reliability leader who can combine SRE discipline, platform operations, and automation-first thinking across the NewLeaf Azure application estate.
Role Summary
- We are hiring a hands-on SRE Lead to own reliability across the full NewLeaf application platform on Azure.
- This role covers customer-facing websites, backend services, APIs, integrations, event-driven workflows, observability, incident management, self-healing automation, on-call operations, and operational reporting.
- The goal is to build a resilient operating model for the entire product ecosystem, not just the underlying cloud infrastructure.
What \"Platform\" Means Here
- Public and internal websites
- Backend services and APIs
- Azure-hosted application workloads
- Integrations and event processing
- Monitoring, alerting, dashboards, and reporting
- Incident response, postmortems, and reliability improvements
- Runbooks, operational standards, and self-healing automation
- On-call rotation design and engineering readiness
Primary Responsibilities
- Define the reliability strategy for the NewLeaf platform and keep it aligned with business priorities.
- Own incident detection, triage, escalation, communication, and post-incident follow-through.
- Build operational dashboards that show service health, user impact, and reliability trends.
- Design self-healing and auto-remediation workflows for recurring failure modes.
- Create and maintain runbooks, playbooks, and a living reliability rulebook.
- Design and support a sustainable on-call rotation model for engineering teams.
- Drive root-cause analysis and convert recurring issues into permanent fixes or automation.
- Partner with application, integration, and platform teams to reduce operational risk across boundaries.
Required Skill Set
- Strong SRE or platform engineering background with direct ownership of production systems.
- Deep incident management experience, including severity classification, incident command, and postmortem rigor.
- Practical observability expertise across logs, metrics, traces, dashboards, alerting, and service health reporting.
- Ability to build or influence automation for detection, triage, remediation, and follow-up tasks.
- Experience shaping on-call practices, escalation paths, and team readiness.
- Comfort operating across application services, integrations, asynchronous processing, and release workflows.
- Clear communication skills for translating technical risk into operational and executive language.
- Ability to coach engineers and raise the maturity of a team without creating unnecessary process overhead.
Azure and Platform Stack Familiarity
- Microsoft Azure operations and governance.
- Azure Monitor, Application Insights, and Log Analytics.
- Azure Service Bus, Event Grid, Event Hub, and timer-driven/background processing patterns.
- Azure Functions, App Service, and other cloud-hosted application runtimes.
- Key Vault, Storage, and identity-aware service configuration.
- CI/CD pipelines and safe deployment practices, including release validation and rollback readiness.
- Infrastructure-as-code tooling such as Bicep or Terraform, where used by the team.
- Kusto Query Language and operational reporting from telemetry data.
Agentic Operations and Automation
- Use agents or automation to assist with analysis, incident summarization, runbook execution, and follow-up work.
- Treat automation as a reliability control, not a novelty layer.
- Apply safety checks, approvals, logging, and rollback paths before any automated remediation is allowed to act on production systems.
- Continuously improve detection and remediation based on real incident patterns and post-incident learning.
Leadership and Collaboration
- Partner with engineering leaders to improve reliability without slowing delivery unnecessarily.
- Lead cross-team incident coordination when production issues span multiple systems.
- Establish practical standards for uptime, alert quality, runbook quality, and operational readiness.
- Help build a culture where reliability is owned by every team, not delegated to a single hero.
Nice-to-Have Experience
- Experience introducing self-healing or auto-remediation in production.
- Experience with executive dashboards and operational scorecards.
- Experience in multi-team or multi-product environments with complex dependencies.
- Background in Azure-native environments, distributed systems, and event-driven architectures.
- Familiarity with AI-assisted operations or agent-based workflows in a safe enterprise setting.
What Success Looks Like
- Incidents are detected faster and resolved faster.
- Recurring reliability issues are systematically eliminated or automated away.
- On-call is sustainable, documented, and well supported.
- Engineers have clear runbooks and confidence during outages.
- Leadership has accurate visibility into platform health and risk.
- Reliability becomes a repeatable operating discipline instead of a hero exercise.