Azure Site Reliability Engineer (SRE) + SQL - SaaS Operations
Job Overview
We are seeking a highly motivated Azure Site Reliability Engineer (SRE) - SaaS Operations to support, maintain, and enhance the reliability of business-critical SaaS applications hosted on Azure cloud environments.
The ideal candidate will have hands‑on experience in Production Support, Application Support, Incident Management, Root Cause Analysis (RCA), Azure Cloud Operations, Monitoring & Observability, Linux Administration, and SQL Troubleshooting. The role requires close collaboration with Engineering, Product, DevOps, Security, and Infrastructure teams to ensure high availability, operational excellence, and continuous service improvement.
Key Responsibilities
- Provide L2/L3 support for production applications and cloud-hosted services.
- Monitor system health, availability, performance, and reliability using enterprise monitoring platforms.
- Handle P1/P2 incidents and drive timely resolution within SLA targets.
- Perform Root Cause Analysis (RCA) for critical incidents and implement preventive measures.
- Support and troubleshoot Azure-based infrastructure and platform services.
- Investigate application, infrastructure, networking, and database-related issues.
- Analyze logs, metrics, dashboards, and alerts to proactively identify service degradation.
- Perform deployment validation, release support, smoke testing, and rollback activities.
- Troubleshoot data issues using SQL queries and database validation techniques.
- Support Linux-based production environments and perform operational troubleshooting.
- Collaborate with Development, Cloud, Infrastructure, and Security teams during major incidents and production releases.
- Participate in on‑call support and operational readiness activities.
- Develop and maintain SOPs, runbooks, knowledge base articles, and operational documentation.
- Drive automation initiatives to improve reliability, efficiency, and operational stability.
Required Skills
Monitoring & Observability
Hands‑on experience with one or more of the following:
- Splunk
- Dynatrace
- Grafana
- ELK Stack
- New Relic
- AppDynamics
Production Support & Operations
- Production Support
- Application Support
- Change Management
- Root Cause Analysis (RCA)
- ITIL Processes
- SLA Management
- On‑call Support
Operating Systems
Database Skills
- SQL
- Query Execution
- Data Validation
- Performance & Issue Troubleshooting
DevOps & Automation (Good to Have)
- Git
- Jenkins
- Docker
- Terraform
- CI/CD Concepts
Additional Skills
- SaaS Application Support
- Cloud Operations
- Release Management
- Deployment Validation
- Performance Monitoring
Preferred Experience
- Experience supporting SaaS‑based applications.
- Experience in Cloud Operations or Site Reliability Engineering (SRE).
- Experience handling customer‑facing production environments.
- Experience working in a 24x7 support model.
- Experience in Banking, Healthcare, Telecom, Retail, Insurance, or Product-based organizations.
Preferred Certifications
- ITIL Foundation
- AWS Cloud Certifications
Ideal Candidate Profile
Professionals currently working as:
- Site Reliability Engineer (SRE)
- Cloud Operations Engineer
- SaaS Operations Engineer
with strong experience in Azure, Incident Management, Monitoring Tools, RCA, Linux, SQL, Production Support, and SaaS Operations.