Bengaluru Tech Infra & IT Site Reliability Engineer - Big Data (4 to 12 Years)

PhonePe Group

Hinoba-an

On-site

PHP 1,500,000 - 2,600,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

PhonePe Limited in the Philippines is seeking a Site Reliability Engineer specialized in Big Data with 7+ years of distributed systems experience. You will manage Linux environments, design automation, and lead on-call incident responses to ensure reliability and security across production clusters.

You will work with Hadoop stack (HDFS, HBase, Kafka, Pinot, etc.), use Puppet/Salt/Ansible, and implement monitoring via ELK, Grafana, and Prometheus.

Qualifications

  • 7+ years of experience managing distributed big data ecosystems.
  • Strong Linux expertise including IP, iptables, and IPsec.
  • Proficiency in scripting languages such as Perl, Golang, or Python.
  • Hands-on experience with Hadoop stack (HDFS, HBase, Airflow, YARN, Ranger, Kafka, Pinot).
  • Familiarity with open-source configuration management and deployment tools such as Puppet, Salt, Chef, or Ansible.
  • Solid understanding of networking, open-source technologies, and related tools.
  • Excellent communication and collaboration skills.
  • DevOps tools: Saltstack, Ansible, Docker, Git.
  • SRE logging and monitoring tools: ELK Stack, Grafana, Prometheus, Open Telemetry.

Responsibilities

  • Manage, maintain, and support incremental changes to Linux/Unix environments.
  • Lead on-call rotations and incident responses, conducting root cause analysis and driving postmortem processes.
  • Design and implement automation systems for managing big data infrastructure, including provisioning, scaling, upgrades, and patching clusters.
  • Troubleshoot and resolve complex production issues while identifying root causes and implementing mitigating strategies.
  • Design and review scalable and reliable system architectures.
  • Collaborate with teams to optimize overall system performance.
  • Enforce security standards across systems and infrastructure.
  • Set technical direction, drive standardization, and operate independently.
  • Ensure availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.
  • Resolve, analyze, and respond to system outages and disruptions and implement measures to prevent similar incidents from recurring.
  • Develop tools and scripts to automate operational processes, reducing manual workload, increasing efficiency and improving system resilience.
  • Monitor and optimize system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
  • Collaborate with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle.
  • Stay informed of industry technology trends and innovations, and actively contribute to the organization's technology communities.
  • Develop and enforce SRE best practices and principles.
  • Align across functional teams on priorities and deliverables.
  • Drive automation to enhance operational efficiency.

Skills

Linux
Scripting languages
Big data stack
Configuration management
Networking basics
DevOps tools
Monitoring/observability

Tools

Puppet
Salt
Chef
Ansible
Docker
Git
ELK Stack
Grafana
Prometheus
OpenTelemetry

Job description

Site Reliability Engineer - Big Data (4 to 12 Years)
  • Full-time

PhonePe Limited (Formerly PhonePe Private Limited) is a technology company that builds digital platforms for Payments, Digital Distribution Services and Financial Services. Headquartered in India, the PhonePe digital payments app was launched in 2016. As of April 2026, PhonePe has over 70 Crore life-till-date registered users and a digital payments acceptance network spread across over 5 Crore merchants. PhonePe’s products and services include Consumer Payments (including Digital Distribution Services), Merchant Payments, Lending and Insurance Distribution services, and New Platforms, which comprise Share.Market (stock broking and mutual funds distribution platform), and Indus Appstore (Android-based mobile app marketplace). Culture:

At PhonePe, we go the extra mile to make sure you can bring your best self to work, Everyday!. And that starts with creating the right environment for you. We empower people and trust them to do the right thing. Here, you own your work from start to finish, right from day one. PhonePe-rs solve complex problems and execute quickly; often building frameworks from scratch. If you’re excited by the idea of building platforms that touch millions, ideating with some of the best minds in the country and executing on your dreams with purpose and speed, join us!

This role is responsible for managing and maintaining complex, distributed big data ecosystems. It ensures the reliability, scalability, and security of large-scale production infrastructure. Key responsibilities include automating processes, optimizing workflows, troubleshooting production issues, and driving system improvements across multiple business verticals.

Roles and Responsibilities:

  • Manage, maintain, and support incremental changes to Linux/Unix environments.
  • Lead on-call rotations and incident responses, conducting root cause analysis and driving postmortem processes.
  • Design and implement automation systems for managing big data infrastructure, including provisioning, scaling, upgrades, and patching clusters.
  • Troubleshoot and resolve complex production issues while identifying root causes and implementing mitigating strategies.
  • Design and review scalable and reliable system architectures.
  • Collaborate with teams to optimize overall system performance.
  • Enforce security standards across systems and infrastructure.
  • Set technical direction, drive standardization, and operate independently.
  • Ensure availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.
  • Resolve, analyze, and respond to system outages and disruptions and implement measures to prevent similar incidents from recurring.
  • Develop tools and scripts to automate operational processes, reducing manual workload, increasing efficiency and improving system resilience.
  • Monitor and optimize system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
  • Collaborate with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle.
  • Stay informed of industry technology trends and innovations, and actively contribute to the organization's technology communities.
  • Develop and enforce SRE best practices and principles.
  • Align across functional teams on priorities and deliverables.
  • Drive automation to enhance operational efficiency.
    • Over 7 years of experience managing and maintaining distributed big data ecosystems.
    • Strong expertise in Linux including IP, Iptables, and IPsec.
    • Proficiency in scripting/programming with languages like Perl, Golang, or Python.
    • Hands‑on experience with the Hadoop stack (HDFS, HBase, Airflow, YARN, Ranger, Kafka, Pinot).
    • Familiarity with open‑source configuration management and deployment tools such as Puppet, Salt, Chef, or Ansible.
    • Solid understanding of networking, open‑source technologies, and related tools.
    • Excellent communication and collaboration skills.
    • DevOps tools: Saltstack, Ansible, docker, Git.
    • SRE Logging and monitoring tools: ELK stack, Grafana, Prometheus, opentsdb, Open Telemetry.

Good to Have:

  • Experience managing infrastructure on public cloud platforms (AWS, Azure, GCP).
  • Experience in designing and reviewing system architectures for scalability and reliability.
  • Experience with observability tools to visualize and alert on system performance.

PhonePe Full Time Employee Benefits (Not applicable for Intern or Contract Roles)

Insurance Benefits - Medical Insurance, Critical Illness Insurance, Accidental Insurance, Life Insurance Wellness Program - Employee Assistance Program, Onsite Medical Center, Emergency Support System Parental Support - Maternity Benefit, Paternity Benefit Program, Adoption Assistance Program, Day-care Support Program Mobility Benefits - Relocation benefits, Transfer Support Policy, Travel Policy Retirement Benefits - Employee PF Contribution, Flexible PF Contribution, Gratuity, NPS, Leave Encashment Other Benefits - Higher Education Assistance, Car Lease, Salary Advance Policy

Our inclusive culture promotes individual expression, creativity, innovation, and achievement and in turn helps us better understand and serve our customers. We see ourselves as a place for intellectual curiosity, ideas and debates, where diverse perspectives lead to deeper understanding and better quality results. PhonePe is an equal opportunity employer and is committed to treating all its employees and job applicants equally; regardless of gender, sexual preference, religion, race, color or disability. If you have a disability or special need that requires assistance or reasonable accommodation, during the application and hiring process, including support for the interview or onboarding process, please fill out this form.

By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply

Job Location
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Bengaluru Tech Infra & IT Site Reliability Engineer - On-Prem (4 to 8 Years)
Bengaluru Tech Infra & IT Site Reliability Engineer - On-Prem (4 to 8 Years)

PhonePe Group • Hinoba-an

On-site
PHP 900,000 - 1,800,000
Insurance Benefits
Wellness Program
Parental Support
+3
Bengaluru Engineering Engineering Manager, Backend
Bengaluru Engineering Engineering Manager, Backend

PhonePe Group • Hinoba-an

On-site
PHP 2,628,000 - 4,600,000
Medical Insurance
Wellness Program
Parental Support
+5
Bengaluru Senior Product Manager Yesterday
Bengaluru Senior Product Manager Yesterday

PhonePe Group • Hinoba-an

On-site
PHP 5,629,000 - 7,505,000
Medical Insurance
Life Insurance
Parental Support
+1
Bengaluru Engineering Head of Engineering - Backend
Bengaluru Engineering Head of Engineering - Backend

PhonePe Group • Hinoba-an

On-site
PHP 3,000,000 - 6,000,000
Medical Insurance
Parental Support
Relocation Benefits
Bengaluru Engineering Engineering Manager - Platforms
Bengaluru Engineering Engineering Manager - Platforms

PhonePe Group • Hinoba-an

On-site
PHP 1,800,000 - 3,200,000
Bengaluru Compliance Manager, Ethics and AC Yesterday
Bengaluru Compliance Manager, Ethics and AC Yesterday

PhonePe Group • Hinoba-an

On-site
PHP 1,200,000 - 2,400,000
Medical Insurance
Parental Support
Relocation Support Policy
Bengaluru Compliance Associate Manager, Technology Risk & Compliance
Bengaluru Compliance Associate Manager, Technology Risk & Compliance

PhonePe Group • Hinoba-an

On-site
PHP 1,200,000 - 2,000,000
Medical Insurance
Critical Illness Insurance
Life Insurance
+4
Bengaluru Compliance ServiceNow Senior Analyst
Bengaluru Compliance ServiceNow Senior Analyst

PhonePe Group • Hinoba-an

On-site
PHP 600,000 - 1,200,000
Medical Insurance
Parental Benefits
Relocation Benefits
+2
Bengaluru Compliance Senior Executive, Technology Risk & Compliance
Bengaluru Compliance Senior Executive, Technology Risk & Compliance

PhonePe Group • Hinoba-an

On-site
PHP 900,000 - 1,200,000
Medical Insurance
Life Insurance
Relocation benefits
+3
Bengaluru Legal Associate Manager, Contract Management
Bengaluru Legal Associate Manager, Contract Management

PhonePe Group • Hinoba-an

On-site
PHP 900,000 - 1,300,000
Medical Insurance
Wellness Program
Parental Support
+1