Site Reliability Engineer 2 (Azure)

Phonepe

Bengaluru

On-site

INR 1,200,000 - 2,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Critical Illness, Accidental, and Life Insurance
Employee Assistance Program
Maternity and Paternity Benefits
Higher Education Assistance

Job summary

Phonepe is seeking a Site Reliability Engineer (SRE) to manage, scale, and ensure the high availability of their core infrastructure. The role requires expertise in cloud services, automation, and complex networking, particularly within a high-volume, mission-critical environment.

Ideal candidates will have deep hands-on experience with Microsoft Azure, strong Linux system skills, and a commitment to drive automations and monitor systems effectively. Benefits include comprehensive insurance, wellness programs, and support for mobility and retirement.

Qualifications

  • Hands-on experience with cloud platforms, especially Microsoft Azure.
  • Expertise in Linux system administration and performance troubleshooting.
  • Strong automation skills with Terraform and experience in Bash scripting.

Responsibilities

  • Manage and ensure high availability of core infrastructure.
  • Implement and maintain monitoring and logging solutions.
  • Conduct proactive capacity planning and manage critical components.

Skills

Microsoft Azure
Linux (Ubuntu)
Terraform
Bash Scripting
Prometheus
MySQL

Tools

Azure Firewall
Docker
Grafana
RabbitMQ
Loki

Job description

Summary

We are seeking a highly motivated and experienced Site Reliability Engineer (SRE) to manage, scale, and ensure the high availability of our core infrastructure. This role involves deep expertise in cloud services, automation, monitoring, and complex networking to support a high‑volume, mission‑critical environment.

Key Responsibilities
  • Cloud & Infrastructure: Configure, maintain, and manage services and packages on Ubuntu Virtual Machines in Azure. Design and manage Azure components for log storage, management, alerting, and monitoring.
  • Networking & Connectivity: Configure and maintain complex network components including Azure Firewall, Route Tables, Virtual Network Gateways, and Express Route. Establish and manage IPsec and Express Route connectivity with external environments. Manage routing, troubleshoot connectivity issues, and support network component migrations with minimal downtime.
  • Automation & IaC: Drive automation for all BAU tasks using Terraform, Saltstack, Ansible, and scripting languages. Write new Terraform code for infrastructure components.
  • Database & Data Management: Set up and manage high‑availability services like Mysql and Aerospike. Implement database replication across regions, manage migrations, and ensure data sync. Handle backups of databases, logs, and configurations.
  • Monitoring & Observability: Implement and manage monitoring (Prometheus, Victoria Metrics, Riemann) and centralized logging (Loki) solutions, with visualization on Grafana. Troubleshoot performance and system issues at the OS, platform, or application level.
  • Security & Compliance: Manage firewalls and integrate platform and VM‑level services with the SOC. Collaborate with Infosec teams to evaluate and fix security vulnerabilities.
  • Capacity & Performance: Conduct proactive capacity planning. Manage critical infrastructure components like Nginx, HA Proxy, Docker, and RMQ.
  • Incident Management & DR: Participate in an on‑call rotation. Structure and lead incident response, Root Cause Analysis (RCA), and post‑mortem creation. Set up and support planning and execution of DR sites and failovers.
Required Technical Expertise
  • Cloud Platform (Microsoft Azure):
    • Core Services: Deep, hands‑on experience with Azure components, including Virtual Machines (Ubuntu/Linux), Azure Storage Accounts, CosmosDB, and Azure Data Explorer.
    • Networking: Expert‑level knowledge in configuring and managing complex Azure networking components: Azure Firewall, Azure Route Tables, Virtual Network Gateways, Azure Express Route, and Azure Private DNS. Must be proficient in setting up and troubleshooting routing using protocols like BGP with on‑prem DCs and managing network component migrations with minimal downtime.
    • Security/Compliance: Experience integrating platform and VM‑level services with the SOC and collaborating with Infosec teams on vulnerability evaluation and remediation.
  • Operating Systems & Scripting:
    • OS: Expert proficiency in Linux environments, specifically Ubuntu/Linux, for system administration, service configuration, and performance troubleshooting at the OS level.
    • High‑Level Language: Deep expertise in at least one high‑level language (Python, Go, or Java) for writing automation, services, and tooling.
    • Shell Scripting: Bash mastery is essential for day‑to‑day operational tasks and automation.
  • Monitoring, Observability & Logging:
    • Monitoring: Extensive experience implementing and maintaining modern monitoring systems such as Prometheus, Victoria Metrics, and Riemann.
    • Logging: Proficiency with centralized log management using Loki for log ingestion, enrichment, lifecycle management, and providing a search/view platform.
    • Visualization: Expertise in creating and managing dashboards for visualization and alerting using Grafana.
  • Configuration Management & IaC:
    • IaC: Mastery of Terraform for writing new component configurations and building automation for BAU (Business As Usual) tasks.
    • Configuration Management: Strong experience with configuration management tools like Saltstack (or Ansible) for automated deployment and configuration of services on VMs.
  • Databases & Data Stores:
    • High‑Availability Data Stores: Hands‑on experience setting up, managing, and scaling high‑availability databases like Mysql and Aerospike.
    • Time‑Series/Search: Familiarity with Elastic Search and time‑series databases like InfluxDB.
    • Replication/DR: Expertise in database replication between different regions, managing database migrations, setting up circular replication, and ensuring data sync during system and network issues.
  • Core Infrastructure Services:
    • Web/Proxy: Expert management of critical infrastructure components like Nginx and HA Proxy, including proxy management, endpoint addition, header configuration, and writing rewrite rules.
    • Messaging/Container: Experience with messaging queues like RMQ (RabbitMQ) and containerization technology like Docker.
    • Networking Services: Deep knowledge of DNS and other core network protocols.
Essential Soft Skills & Qualifications
  • Ownership and Accountability: Proactive approach to identifying and solving infrastructure challenges before they impact service availability.
  • Communication: Excellent written and verbal skills for documenting procedures, creating runbooks, and communicating with technical and non‑technical stakeholders.
  • Mentorship: (For senior roles) Ability to mentor junior engineers and promote SRE best practices across the organization.
  • SLO/SLA Management: Experience defining, monitoring, and meeting Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical services.
  • Toil Reduction: Commitment to measuring and actively reducing operational toil through automation (e.g., using SRE's Toil Reduction framework).
  • Cost Optimization: Experience identifying and implementing cloud resource optimization and cost‑saving measures within the Azure environment.
Benefits
  • Insurance Benefits: Medical, Critical Illness, Accidental, and Life Insurance.
  • Wellness Program: Employee Assistance Program, Onsite Medical Center, and Emergency Support System.
  • Parental Support: Maternity and Paternity Benefits, Adoption Assistance Program, and Day‑care Support Program.
  • Mobility Benefits: Relocation benefits, Transfer Support Policy, and Travel Policy.
  • Retirement Benefits: Employee PF Contribution, Flexible PF Contribution, Gratuity, NPS, and Leave Encashment.
  • Other Benefits: Higher Education Assistance, Car Lease, and Salary Advance Policy.
Equal Opportunity Employer

PhonePe is an equal opportunity employer and is committed to treating all its employees and job applicants equally; regardless of gender, sexual preference, religion, race, color or disability. If you have a disability or special need that requires assistance or reasonable accommodation during the application and hiring process, please let us know.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Phonepe • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Medical Insurance
Wellness Program
Parental Support
+2
Site Reliability Engineer (4 to 8 Years)
Site Reliability Engineer (4 to 8 Years)

PhonePe • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Medical Insurance
Critical Illness Insurance
Accidental Insurance
+6
Site Reliability Engineer 2 Years
Site Reliability Engineer 2 Years

PhonePe • Bengaluru

On-site
INR 900,000 - 1,500,000
Insurance Benefits
Wellness Program
Parental Support
+3
Site Reliability Engineer (4 to 8 Years)
Site Reliability Engineer (4 to 8 Years)

PhonePe • Bengaluru

On-site
INR 1,650,000 - 2,100,000
Insurance
Wellness Program
Parental support
+3
Site Reliability Engineer (2+ Years)
Site Reliability Engineer (2+ Years)

PhonePe • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Medical Insurance
Critical Illness Insurance
Accidental Insurance
+16
Site Reliability Engineer (4+ YOE)
Site Reliability Engineer (4+ YOE)

PhonePe • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Medical Insurance
Critical Illness Insurance
Accidental Insurance
+19
Service Delivery Engineer, SRE
Service Delivery Engineer, SRE

PhonePe • Bengaluru

On-site
INR 800,000 - 1,500,000
Medical, Critical Illness, Accidental, Life Insurance
Employee Assistance Program
Maternity and Paternity Benefits
+1
Bengaluru - Salarpuria Softzone (SSZ) Tech Infra & IT Site Reliability Engineer (4 to 8 Years)
Bengaluru - Salarpuria Softzone (SSZ) Tech Infra & IT Site Reliability Engineer (4 to 8 Years)

PhonePe Group • Bengaluru

On-site
INR 1,200,000 - 2,200,000
Medical Insurance
Critical Illness Insurance
Relocation Benefits
+1
Site Reliability Engineer
Site Reliability Engineer

Plume • Hyderabad

On-site
INR 2,500,000 - 5,200,000
Lead SRE & Support Engineer
Lead SRE & Support Engineer

Providence Global Center • Hyderabad

On-site
INR 3,500,000 - 5,500,000
Competitive Pay
Supportive Reporting Relation