Site Reliability Engineer

Finthrive

Gurugram District

Hybrid

INR 1,800,000 - 2,400,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Finthrive is seeking an experienced Site Reliability Engineer / Cloud Engineer in Gurugram to own cloud deployments with a focus on reliability, automation, and observability. You will lead incident responses, perform RCAs, and implement preventive measures to increase availability of cloud-hosted applications.

The role emphasizes Azure App Services, ASEv3, Front Door, AGW tuning, and IaC using Terraform, Bicep, and ARM templates, along with AI-assisted tooling for scripting and efficient

Qualifications

  • Strong hands-on experience with Microsoft Azure services and cloud architectures.
  • Proficient in building and operating scalable, highly available cloud deployments.
  • Experience in automation, IaC, and incident response processes.

Responsibilities

  • Manage production cloud environments with high availability and reliability.
  • Lead incident response, perform RCAs, and implement preventive measures.
  • Enhance observability and tuning to reduce alert fatigue and improve performance.

Skills

Azure Cloud
CI/CD
Observability
Automation
Incident Management
Networking
Terraform
Azure Functions
KQL
DevOps

Education

Bachelor's in CS/Engineering

Tools

ARM Templates
Azure Bicep
Azure Automation
GitHub Copilot
Front Door
Application Gateway

Job description

Role & responsibilities

Site Reliability Engineer / Cloud Engineer

SRE & Reliability Engineering

  • Managed production environments ensuring high availability and reliability of cloud-hosted applications
  • Led incident response, performed deep root cause analysis, and implemented preventive measures to reduce recurrence
  • Improved system resilience through proactive monitoring and performance tuning strategies

Azure Application & Platform Engineering

  • Designed and supported application architectures using:
  • Azure App Services and App Service Plans
  • Azure App Service Environment v3 (ASEv3) for isolated, high-scale workloads
  • Azure Application Gateway (WAF-enabled) for L7 traffic management
  • Azure Front Door for global traffic routing and failover
  • Implemented secure and scalable cloud networking patterns, optimizing latency and throughput

Automation & Toil Reduction

  • Identified repetitive operational tasks and reduced manual effort through automation-first solutions
  • Developed automation using:
  • Terraform / Bicep / ARM templates
  • Azure Automation (Hybrid Workers)
  • Azure Functions for event-driven workflows
  • Leveraged AI-assisted tools (GitHub Copilot, Copilot) to accelerate scripting and automation development, while ensuring strict validation for enterprise use

Observability & Monitoring

  • Built and enhanced observability using:
  • Azure Monitor, Application Insights, Log Analytics
  • Created KQL-based queries and dashboards for proactive issue detection
  • Reduced false alerts by optimizing alert thresholds and improving signal quality

Performance & System Optimization

  • Analyzed application performance across distributed systems to identify bottlenecks
  • Implemented improvements through:
  • Scaling strategies (horizontal & vertical)
  • Network optimization (AGW / Front Door tuning)
  • Backend service improvements

Collaboration & Engineering Enablement

  • Partnered with SRE, CloudOps, and development teams to design resilient systems
  • Contributed to runbooks, documentation, and operational standards
  • Enabled engineering teams by improving platform reliability and deployment pipelines

Key Achievements

  • Reduced manual operational effort by X% through automation initiatives
  • Improved system availability to 99.X% by strengthening monitoring and failure handling mechanisms
  • Decreased incident resolution time by X% via enhanced observability and streamlined runbooks
  • Optimized application performance using Front Door and AGW tuning, reducing latency by X%

Preferred candidate profile

Cloud & Platform Engineering

  • Microsoft Azure (Preferred)
  • Understanding and experience in developing Azure function Apps, Azure logic Apps
  • Understanding of event triggers, event hub, service bus.
  • Azure Landing zones, Azure Cloud Adoption Framework, Azure Well Architectured Framework
  • Application Hosting: App Services, App Service Plans, ASEv3
  • Networking: Azure Application Gateway (AGW), Azure Front Door, VNet, NSGs, Load Balancing
  • Cloud Architecture: High Availability, Fault Tolerance, Scalability Patterns

Incident Management and RCA

  • Incident Management, P1 troubleshooting, Change Management
  • Experienced in leading RCA and representing on the weekly call
  • SLA / SLO / Error Budget concepts
  • System Performance Optimization & Capacity Planning
  • Toil Reduction through Automation

Infrastructure as Code & Automation

  • Terraform, Azure Bicep, ARM Templates
  • Azure Automation (Hybrid Workers)
  • Azure Functions (Serverless automation)
  • API-based automation and orchestration

Observability & Monitoring

  • Azure Monitor, Log Analytics Workspace, Grafana, Site 24x7 (or similar SaaS based synthetic monitoring tool)
  • Application Insights
  • KQL (Kusto Query Language)
  • Alert tuning and signal-to-noise optimization

AI-Enabled Productivity (Not as Skill)

  • Leveraging GitHub Copilot / Microsoft Copilot for:
  • Code acceleration and script generation
  • Automation development support
  • Troubleshooting and log analysis assistance
  • Proven track record of workforce optimization leveraging AI tools.
  • Applying validation frameworks to ensure secure, accurate, and production-grade outputs

DevOps & Integration

  • CI/CD using Azure DevOps
  • Deep understanding on version control
  • API integrations (REST, Postman, SoapUI)
  • Source control and release management
  • Bachelors Degree in Computer Science / Engineering or related field

Preferred/Additional Experience

  • Experience with microservices and distributed architectures
  • Exposure to low-code automation platforms
  • Working knowledge of AWS cloud services

Preferred/Additional Certifications

  • AZ-104 Azure Administrator
  • AZ-700 Designing and Implementing Microsoft Azure Networking Solutions
  • AZ-400 Microsoft Certified: DevOps Engineer Expert
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (Azure) - S
Senior Site Reliability Engineer (Azure) - S

Tata Consultancy Services • Kolkata District, Chennai District, Bengaluru

On-site
INR 2,400,000 - 4,200,000
Site Reliability Engineer
Site Reliability Engineer

PwC India • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Azure Platform Site Reliability Engineer
Azure Platform Site Reliability Engineer

Foss United • India

On-site
INR 2,800,000 - 4,600,000
Senior Site Reliability Engineering (Azure Cloud)
Senior Site Reliability Engineering (Azure Cloud)

Cvent • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Azure Cloud SRE
Azure Cloud SRE

Intellics Global Services • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Embarkgcc Services • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Azure Site Reliability Engineer (SRE) + SQL - SaaS Operations
Azure Site Reliability Engineer (SRE) + SQL - SaaS Operations

Zensar Technologies • Pune District

On-site
INR 2,000,000 - 2,800,000
Azure Senior Site Reliability Engineer
Azure Senior Site Reliability Engineer

LTM • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Azure Cloud Engineer
Azure Cloud Engineer

Jobtailor • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Director of DevOps ,SRE and Infrastructure | (Azure) |12+ Years | Ahmedabad Location
Director of DevOps ,SRE and Infrastructure | (Azure) |12+ Years | Ahmedabad Location

Vagaro • Ahmedabad District

On-site
INR 4,000,000 - 6,000,000
5-day work week
Flexible schedule
Annual bonus
+3