Site Reliability Engineer 3
Summary
We are seeking a highly motivated and experienced Site Reliability Engineer (SRE) with 7 to 12 years of experience to manage, scale, and ensure the high availability of our core infrastructure. This role involves deep expertise in cloud services, automation, monitoring, and complex networking to support a high-volume, mission‑critical environment.
Key Responsibilities
- Cloud & Infrastructure: Configure, maintain, and manage services and packages on Ubuntu Virtual Machines in Azure; design and manage Azure components for log storage, management, alerting, and monitoring.
- Networking & Connectivity: Configure and maintain complex network components including Azure Firewall, Route Tables, Virtual Network Gateways, Express Route; establish and manage IPsec and Express Route connectivity with external environments; manage routing, troubleshoot connectivity issues, and support network component migrations with minimal downtime.
- Automation & IaC: Drive automation for BAU tasks using Terraform, SaltStack, Ansible, and scripting languages; write new Terraform code for infrastructure components.
- Database & Data Management: Set up and manage high‑availability services like MySQL and Aerospike; implement database replication across regions, manage migrations, ensure data sync; handle backups of databases, logs, and configurations.
- Monitoring & Observability: Implement and manage monitoring (Prometheus, Victoria Metrics, Riemann) and centralized logging (Loki) solutions with visualization on Grafana; troubleshoot performance and system issues at OS, platform, or application level.
- Security & Compliance: Manage firewalls and integrate platform and VM‑level services with the SOC; collaborate with Infosec teams to evaluate and fix security vulnerabilities.
- Capacity & Performance: Conduct proactive capacity planning; manage critical infrastructure components like Nginx, HA Proxy, Docker, and RMQ.
- Incident Management & DR: Participate in an on‑call rotation; structure and lead incident response, root‑cause analysis, and post‑mortem creation; set up and support planning and execution of DR sites and failovers.
Required Technical Expertise
- Cloud Platform (Microsoft Azure): Deep hands‑on experience with Azure Virtual Machines (Ubuntu/Linux), Azure Storage Accounts, CosmosDB, and Azure Data Explorer (ADX).
- Networking: Expert knowledge in configuring and managing Azure Firewall, Azure Route Tables, Virtual Network Gateways, Azure Express Route, Azure Private DNS, BGP routing with on‑prem DCs, and managing network component migrations.
- Security / Compliance: Experience integrating platform and VM‑level services with SOC and collaborating with Infosec teams.
- Operating Systems & Scripting: Expert proficiency in Linux (Ubuntu/Linux); deep expertise in one high‑level language (Python, Go, or Java); mastery of Bash scripting.
- Monitoring, Observability & Logging: Extensive experience with Prometheus, Victoria Metrics, Riemann, Loki, and Grafana dashboards.
- Infrastructure as Code: Mastery of Terraform; strong experience with SaltStack or Ansible.
- Databases & Data Stores: Experience setting up, managing, and scaling MySQL and Aerospike; familiarity with Elasticsearch, InfluxDB, database replication, and DR.
- Core Infrastructure Services: Management of Nginx, HA Proxy, RabbitMQ, Docker; deep knowledge of DNS and core network protocols.
Soft Skills & Qualifications
- Ownership and accountability with a proactive approach to infrastructure challenges.
- Excellent written and verbal communication for documenting procedures, runbooks, and stakeholder communication.
- Mentorship experience for senior roles.
- Experience defining and monitoring SLOs and SLIs.
- Commitment to toil reduction and cost optimization in Azure.
Benefits
- Medical, Critical Illness, Accidental, Life Insurance.
- Wellness Program: Employee Assistance Program, On‑site Medical Center, Emergency Support System.
- Parental Support: Maternity, Paternity, Adoption Assistance, Day‑care Support.
- Mobility: Relocation benefits, Transfer Support Policy, Travel Policy.
- Retirement: Employee PF Contribution, Flexible PF Contribution, Gratuity, NPS, Leave Encashment.
- Other: Higher Education Assistance, Car Lease, Salary Advance Policy.
Equal Opportunity Employer
PhonePe is an equal opportunity employer and is committed to treating all its employees and job applicants equally; regardless of gender, sexual preference, religion, race, color or disability. If you have a disability or special need that requires assistance or reasonable accommodation, please fill out the form provided in the application process.