Site Reliability Engineer

Phonepe

Bengaluru

On-site

INR 1,500,000 - 2,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical Insurance
Wellness Program
Parental Support
Retirement Benefits
Higher Education Assistance

Job summary

Phonepe is seeking an experienced Site Reliability Engineer (SRE) in Bengaluru to manage and scale core infrastructure. Ideal candidates have 7 to 12 years of experience, expertise in cloud services, and proficiency in automation and monitoring tools.

Responsibilities include managing Azure services, driving automation, and ensuring high availability. Candidates must demonstrate strong networking skills and proficiency in Linux, as well as have excellent communication abilities.

Qualifications

  • 7 to 12 years of experience as a Site Reliability Engineer.
  • Deep expertise in cloud services and automation.
  • Excellent communication skills for documentation and stakeholder interactions.

Responsibilities

  • Manage and ensure high availability of core infrastructure.
  • Drive automation for BAU tasks using Terraform and scripting languages.
  • Participate in incident response and disaster recovery planning.

Skills

Azure Virtual Machines
Terraform
Linux
Networking
Prometheus
MySQL
Bash scripting
Security compliance

Job description

Site Reliability Engineer 3

Summary

We are seeking a highly motivated and experienced Site Reliability Engineer (SRE) with 7 to 12 years of experience to manage, scale, and ensure the high availability of our core infrastructure. This role involves deep expertise in cloud services, automation, monitoring, and complex networking to support a high-volume, mission‑critical environment.

Key Responsibilities
  • Cloud & Infrastructure: Configure, maintain, and manage services and packages on Ubuntu Virtual Machines in Azure; design and manage Azure components for log storage, management, alerting, and monitoring.
  • Networking & Connectivity: Configure and maintain complex network components including Azure Firewall, Route Tables, Virtual Network Gateways, Express Route; establish and manage IPsec and Express Route connectivity with external environments; manage routing, troubleshoot connectivity issues, and support network component migrations with minimal downtime.
  • Automation & IaC: Drive automation for BAU tasks using Terraform, SaltStack, Ansible, and scripting languages; write new Terraform code for infrastructure components.
  • Database & Data Management: Set up and manage high‑availability services like MySQL and Aerospike; implement database replication across regions, manage migrations, ensure data sync; handle backups of databases, logs, and configurations.
  • Monitoring & Observability: Implement and manage monitoring (Prometheus, Victoria Metrics, Riemann) and centralized logging (Loki) solutions with visualization on Grafana; troubleshoot performance and system issues at OS, platform, or application level.
  • Security & Compliance: Manage firewalls and integrate platform and VM‑level services with the SOC; collaborate with Infosec teams to evaluate and fix security vulnerabilities.
  • Capacity & Performance: Conduct proactive capacity planning; manage critical infrastructure components like Nginx, HA Proxy, Docker, and RMQ.
  • Incident Management & DR: Participate in an on‑call rotation; structure and lead incident response, root‑cause analysis, and post‑mortem creation; set up and support planning and execution of DR sites and failovers.
Required Technical Expertise
  • Cloud Platform (Microsoft Azure): Deep hands‑on experience with Azure Virtual Machines (Ubuntu/Linux), Azure Storage Accounts, CosmosDB, and Azure Data Explorer (ADX).
  • Networking: Expert knowledge in configuring and managing Azure Firewall, Azure Route Tables, Virtual Network Gateways, Azure Express Route, Azure Private DNS, BGP routing with on‑prem DCs, and managing network component migrations.
  • Security / Compliance: Experience integrating platform and VM‑level services with SOC and collaborating with Infosec teams.
  • Operating Systems & Scripting: Expert proficiency in Linux (Ubuntu/Linux); deep expertise in one high‑level language (Python, Go, or Java); mastery of Bash scripting.
  • Monitoring, Observability & Logging: Extensive experience with Prometheus, Victoria Metrics, Riemann, Loki, and Grafana dashboards.
  • Infrastructure as Code: Mastery of Terraform; strong experience with SaltStack or Ansible.
  • Databases & Data Stores: Experience setting up, managing, and scaling MySQL and Aerospike; familiarity with Elasticsearch, InfluxDB, database replication, and DR.
  • Core Infrastructure Services: Management of Nginx, HA Proxy, RabbitMQ, Docker; deep knowledge of DNS and core network protocols.
Soft Skills & Qualifications
  • Ownership and accountability with a proactive approach to infrastructure challenges.
  • Excellent written and verbal communication for documenting procedures, runbooks, and stakeholder communication.
  • Mentorship experience for senior roles.
  • Experience defining and monitoring SLOs and SLIs.
  • Commitment to toil reduction and cost optimization in Azure.
Benefits
  • Medical, Critical Illness, Accidental, Life Insurance.
  • Wellness Program: Employee Assistance Program, On‑site Medical Center, Emergency Support System.
  • Parental Support: Maternity, Paternity, Adoption Assistance, Day‑care Support.
  • Mobility: Relocation benefits, Transfer Support Policy, Travel Policy.
  • Retirement: Employee PF Contribution, Flexible PF Contribution, Gratuity, NPS, Leave Encashment.
  • Other: Higher Education Assistance, Car Lease, Salary Advance Policy.
Equal Opportunity Employer

PhonePe is an equal opportunity employer and is committed to treating all its employees and job applicants equally; regardless of gender, sexual preference, religion, race, color or disability. If you have a disability or special need that requires assistance or reasonable accommodation, please fill out the form provided in the application process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer 2 (Azure)
Site Reliability Engineer 2 (Azure)

Phonepe • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Medical, Critical Illness, Accidental, and Life Insurance
Employee Assistance Program
Maternity and Paternity Benefits
+1
Site Reliability Engineer (4 to 8 Years)
Site Reliability Engineer (4 to 8 Years)

PhonePe • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Medical Insurance
Critical Illness Insurance
Accidental Insurance
+6
Site Reliability Engineer 2 Years
Site Reliability Engineer 2 Years

PhonePe • Bengaluru

On-site
INR 900,000 - 1,500,000
Insurance Benefits
Wellness Program
Parental Support
+3
Site Reliability Engineer (4 to 8 Years)
Site Reliability Engineer (4 to 8 Years)

PhonePe • Bengaluru

On-site
INR 1,650,000 - 2,100,000
Insurance
Wellness Program
Parental support
+3
Site Reliability Engineer (2+ Years)
Site Reliability Engineer (2+ Years)

PhonePe • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Medical Insurance
Critical Illness Insurance
Accidental Insurance
+16
Site Reliability Engineer
Site Reliability Engineer

Linuxcareers • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Medical Insurance
Maternity and Paternity benefits
Higher Education Assistance
+1
Site Reliability Engineer (4+ YOE)
Site Reliability Engineer (4+ YOE)

PhonePe • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Medical Insurance
Critical Illness Insurance
Accidental Insurance
+19
Bengaluru - Salarpuria Softzone (SSZ) Tech Infra & IT Site Reliability Engineer (4 to 8 Years)
Bengaluru - Salarpuria Softzone (SSZ) Tech Infra & IT Site Reliability Engineer (4 to 8 Years)

PhonePe Group • Bengaluru

On-site
INR 1,200,000 - 2,200,000
Medical Insurance
Critical Illness Insurance
Relocation Benefits
+1
Site Reliability Engineer
Site Reliability Engineer

Plume • Hyderabad

On-site
INR 2,500,000 - 5,200,000
Site Reliability Engineer
Site Reliability Engineer

MishiPay • Bengaluru

On-site
INR 2,500,000 - 3,800,000