Site Reliability Engineer - AWS (7 to 12 Years)

PhonePe

Bengaluru

On-site

INR 3,500,000 - 5,500,000

Full time

47 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical Insurance
Critical Illness Insurance
Accidental Insurance
Life Insurance
Wellness Program
Employee Assistance Program
Onsite Medical Center
Parental Benefits
Relocation Benefits
Transfer Support Policy
Travel Policy
PF Contribution
Gratuity
NPS
Leave Encashment
Higher Education Assistance
Car Lease
Salary Advance Policy

Job summary

PhonePe is seeking a Site Reliability Engineer (SRE) with 7+ years of experience to manage and scale our core infrastructure. The role focuses on AWS-based cloud architecture, Linux/ RHEL administration, high availability, and security compliance to support a high-demand environment.

You will lead advanced network setups, containerization, automation, and disaster recovery planning across regions. This position offers a challenging, growth-oriented path within PhonePe's engineering team.

Qualifications

  • 7+ years as a Senior SRE, Cloud Architect, or Lead Systems Engineer.
  • Advanced proficiency in Linux/RHEL administration, troubleshooting, kernel tuning, and security hardening.
  • Deep architectural expertise within the AWS ecosystem.

Responsibilities

  • Lead Linux/RHEL administration with kernel tuning and security hardening for high-volume environments.
  • Build and scale highly available AWS environments to handle massive traffic.
  • Architect infrastructure with HA and security standards; drive compliance best practices across systems.
  • Design and maintain network connectivity with AWS Direct Connect, Transit Gateways, VPNs, and BGP.
  • Plan large-scale migrations, upgrades, and database sharding with minimal production disruption.
  • Lead capacity planning for growth and traffic spikes.
  • Define and promote IaC and configuration management practices; build self-healing automation.
  • Design cross-region database replication, HA scaling, and disaster recovery for MySQL/Aerospike.
  • Establish reliability metrics and lead incident response with RCAs to reduce debt.
  • Document system architecture, disaster recovery plans, and SOPs; mentor teammates.

Skills

Linux/RHEL admin
AWS cloud architecture
Network engineering (BGP)
MySQL administration
Load balancers (Nginx/HAProxy)
Containerization (Docker/Podman)
Infrastructure as Code
Security hardening/compliance
Technical communication

Tools

Docker
Podman
Nginx
HAProxy
AWS Direct Connect
Transit Gateway
BGP tooling

Job description

PhonePe Limited (Formerly PhonePe Private Limited) is a technology company that builds digital platforms for Payments, Digital Distribution Services and Financial Services. Headquartered in India, the PhonePe digital payments app was launched in 2016. As of April 2026, PhonePe has over 70 Crore life-till-date registered users and a digital payments acceptance network spread across over 5 Crore merchants. PhonePe’s products and services include Consumer Payments (including Digital Distribution Services), Merchant Payments, Lending and Insurance Distribution services, and New Platforms, which comprise Share.Market (stock broking and mutual funds distribution platform), and Indus Appstore (Android-based mobile app marketplace).

Culture

At PhonePe, we go the extra mile to make sure you can bring your best self to work, Everyday! And that starts with creating the right environment for you. We empower people and trust them to do the right thing. Here, you own your work from start to finish, right from day one. PhonePe-rs solve complex problems and execute quickly; often building frameworks from scratch. If you’re excited by the idea of building platforms that touch millions, ideating with some of the best minds in the country and executing on your dreams with purpose and speed, join us!

Job Description

We are seeking a highly motivated Site Reliability Engineer (SRE) with more than 7 years of experience to manage, scale, and ensure the high availability of our core infrastructure. This role is designed for experts specialized in AWS. You will leverage a profound background in Linux (specifically RHEL) to drive deep-level cloud architecture, automation, complex networking, and security compliance, supporting a high-volume, mission-critical environment that demands exceptional uptime and resilience.

Roles And Responsibilities
  • Lead advanced Linux/RHEL administration, focusing on OS troubleshooting, kernel-level performance tuning, and executing security hardening for high-volume, mission-critical environments.
  • Build and scale highly available AWS environments that can handle massive traffic without failing. Drive the long-term strategy for our servers, data storage, and system monitoring.
  • Architect infrastructure with a strong emphasis on High Availability (HA) and robust Information Security. Set and drive the adoption of security standards and compliance best practices across the infrastructure.
  • Design and maintain advanced network connectivity and routing using AWS Direct Connect, Transit Gateways, VPNs, and BGP
  • Plan, strategize, and oversee large-scale infrastructure migrations, component upgrades, and database sharding while ensuring minimal to zero disruption to production
  • Lead capacity planning for organic growth and traffic spikes.
  • Define the best practices and tools we use for Infrastructure as Code (IaC) and configuration management. Continuously reduce manual toil by designing automated, self-healing systems.
  • Design and implement cross-region database replication, HA scaling strategies, and comprehensive disaster recovery plans for data stores like MySQL and Aerospike.
  • Establish and track our reliability metrics. Lead the response during production outages, and use post-mortems (RCAs) to fix underlying architectural flaws and reduce technical debt
  • Establish documentation standards for system architecture, disaster recovery plans, and SOPs. Mentor fellow team members and foster a collaborative culture of reliability and continuous learning.
Qualifications
Minimum Requirements
  • Experience: 7+ years as a Senior SRE, Cloud Architect, or Lead Systems Engineer.
  • Advanced proficiency in Linux/RHEL administration, including hands-on troubleshooting, kernel-level performance tuning, and OS security hardening.
  • Deep, architectural-level expertise exclusively within the AWS ecosystem.
  • Strong understanding of BGP routing, complex TCP/IP troubleshooting, and large-scale network design.
  • Understanding of MySQL/MariaDB/Percona or any other relational database concepts and administration
  • Proven experience defining and tracking site reliability, alongside a strong passion for automation.
  • Proven ability to scale and optimize load balancers (Nginx, HAProxy) to handle massive web traffic.
  • Hands-on experience orchestrating and scaling containerization technologies (Docker, Podman, or equivalent).
  • Ability to effectively communicate complex technical strategies and architectural decisions to internal teams and external stakeholders.
Preferred Qualifications (A Plus)
  • Experience optimizing cloud costs - finding wasted resources and cutting AWS bills without sacrificing performance.
  • Knowledge of NoSQL databases like Aerospike
  • Experience with distributed event streaming platforms like Kafka.
  • Experience predicting future infrastructure needs and planning capacity for long-term growth.
Additional Information
PhonePe Full Time Employee Benefits (Not applicable for Intern or Contract Roles)
  • Insurance Benefits - Medical Insurance, Critical Illness Insurance, Accidental Insurance, Life Insurance
  • Wellness Program - Employee Assistance Program, Onsite Medical Center, Emergency Support System
  • Parental Support - Maternity Benefit, Paternity Benefit Program, Adoption Assistance Program, Day-care Support Program
  • Mobility Benefits - Relocation benefits, Transfer Support Policy, Travel Policy
  • Retirement Benefits - Employee PF Contribution, Flexible PF Contribution, Gratuity, NPS, Leave Encashment
  • Other Benefits - Higher Education Assistance, Car Lease, Salary Advance Policy

Our inclusive culture promotes individual expression, creativity, innovation, and achievement and in turn helps us better understand and serve our customers. We see ourselves as a place for intellectual curiosity, ideas and debates, where diverse perspectives lead to deeper understanding and better quality results. PhonePe is an equal opportunity employer and is committed to treating all its employees and job applicants equally; regardless of gender, sexual preference, religion, race, color or disability.

If you have a disability or special need that requires assistance or reasonable accommodation, during the application and hiring process, including support for the interview or onboarding process, please fill out this form.

Read more about PhonePe on our blog.

Life at PhonePe

PhonePe in the news

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - On-Prem (4 to 8 Years)
Site Reliability Engineer - On-Prem (4 to 8 Years)

PhonePe • Bengaluru

On-site
INR 1,800,000 - 2,600,000
Insurance Benefits - Medical, Critical
Wellness Program
Parental Support
+3
Site Reliability Engineer - On-Prem (4 to 12 Years)
Site Reliability Engineer - On-Prem (4 to 12 Years)

PhonePe • Bengaluru

On-site
INR 1,200,000 - 1,600,000
Medical Insurance
Critical Illness Insurance
Accidental Insurance
+3
Site Reliability Engineer 2 Years
Site Reliability Engineer 2 Years

PhonePe • Bengaluru

On-site
INR 900,000 - 1,500,000
Insurance Benefits
Wellness Program
Parental Support
+3
Site Reliability Engineer (4 to 8 Years)
Site Reliability Engineer (4 to 8 Years)

PhonePe • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Medical Insurance
Critical Illness Insurance
Accidental Insurance
+6
Site Reliability Engineer - On-Prem (2 to 8 Years)
Site Reliability Engineer - On-Prem (2 to 8 Years)

PhonePe • Bengaluru

On-site
INR 1,100,000 - 1,400,000
Medical Insurance
Critical Illness Insurance
Accidental Insurance
+4
Site Reliability Engineer (4+ YOE)
Site Reliability Engineer (4+ YOE)

PhonePe • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Medical Insurance
Critical Illness Insurance
Accidental Insurance
+19
Site Reliability Engineer 2
Site Reliability Engineer 2

PhonePe • Bengaluru

On-site
INR 1,000,000 - 2,000,000
Medical Insurance
Wellness Program
Maternity Benefit
+2
Head of Engineering - Backend
Head of Engineering - Backend

PhonePe • Bengaluru

On-site
INR 4,000,000 - 7,500,000
Medical & insurance benefits
Parental & family support
Relocation & travel policy
+2
Site Reliability Engineer 2 (Azure)
Site Reliability Engineer 2 (Azure)

Phonepe • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Medical, Critical Illness, Accidental, and Life Insurance
Employee Assistance Program
Maternity and Paternity Benefits
+1
Service Delivery Engineer, SRE
Service Delivery Engineer, SRE

PhonePe • Bengaluru

On-site
INR 800,000 - 1,500,000
Medical, Critical Illness, Accidental, Life Insurance
Employee Assistance Program
Maternity and Paternity Benefits
+1