Senior Big Data SRE: Automation & Reliability

PhonePe Group

Hinoba-an

On-site

PHP 1,500,000 - 2,600,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

PhonePe Limited in the Philippines is seeking a Site Reliability Engineer specialized in Big Data with 7+ years of distributed systems experience. You will manage Linux environments, design automation, and lead on-call incident responses to ensure reliability and security across production clusters.

You will work with Hadoop stack (HDFS, HBase, Kafka, Pinot, etc.), use Puppet/Salt/Ansible, and implement monitoring via ELK, Grafana, and Prometheus.

Qualifications

  • 7+ years of experience managing distributed big data ecosystems.
  • Strong Linux expertise including IP, iptables, and IPsec.
  • Proficiency in scripting languages such as Perl, Golang, or Python.
  • Hands-on experience with Hadoop stack (HDFS, HBase, Airflow, YARN, Ranger, Kafka, Pinot).
  • Familiarity with open-source configuration management and deployment tools such as Puppet, Salt, Chef, or Ansible.
  • Solid understanding of networking, open-source technologies, and related tools.
  • Excellent communication and collaboration skills.
  • DevOps tools: Saltstack, Ansible, Docker, Git.
  • SRE logging and monitoring tools: ELK Stack, Grafana, Prometheus, Open Telemetry.

Responsibilities

  • Manage, maintain, and support incremental changes to Linux/Unix environments.
  • Lead on-call rotations and incident responses, conducting root cause analysis and driving postmortem processes.
  • Design and implement automation systems for managing big data infrastructure, including provisioning, scaling, upgrades, and patching clusters.
  • Troubleshoot and resolve complex production issues while identifying root causes and implementing mitigating strategies.
  • Design and review scalable and reliable system architectures.
  • Collaborate with teams to optimize overall system performance.
  • Enforce security standards across systems and infrastructure.
  • Set technical direction, drive standardization, and operate independently.
  • Ensure availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.
  • Resolve, analyze, and respond to system outages and disruptions and implement measures to prevent similar incidents from recurring.
  • Develop tools and scripts to automate operational processes, reducing manual workload, increasing efficiency and improving system resilience.
  • Monitor and optimize system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
  • Collaborate with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle.
  • Stay informed of industry technology trends and innovations, and actively contribute to the organization's technology communities.
  • Develop and enforce SRE best practices and principles.
  • Align across functional teams on priorities and deliverables.
  • Drive automation to enhance operational efficiency.

Skills

Linux
Scripting languages
Big data stack
Configuration management
Networking basics
DevOps tools
Monitoring/observability

Tools

Puppet
Salt
Chef
Ansible
Docker
Git
ELK Stack
Grafana
Prometheus
OpenTelemetry

Job description

PhonePe Limited in the Philippines is seeking a Site Reliability Engineer specialized in Big Data with 7+ years of distributed systems experience. You will manage Linux environments, design automation, and lead on-call incident responses to ensure reliability and security across production clusters.

You will work with Hadoop stack (HDFS, HBase, Kafka, Pinot, etc.), use Puppet/Salt/Ansible, and implement monitoring via ELK, Grafana, and Prometheus.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Application Engineer, DevOps & SRE for Reliability
Senior Application Engineer, DevOps & SRE for Reliability

Home Credit Philippines • Taguig

On-site
PHP 900,000 - 1,500,000
Permanent dayshift
Performance bonus
HMO coverage
+4
Senior SRE: Observability, Cloud Automation & Reliability
Senior SRE: Observability, Cloud Automation & Reliability

Omilia Natural Language Solutions Ua Ltd • Philippines

On-site
PHP 1,000,000 - 1,800,000
Fixed compensation
Long-term vacation
Career development courses
+1
Senior DevOps/SRE Engineer - Reliability & Automation Lead
Senior DevOps/SRE Engineer - Reliability & Automation Lead

Home Credit Philippines • Philippines

On-site
PHP 900,000 - 1,700,000
Permanent dayshift schedule
Up to 20% performance bonus
HMO on Day 1 and dependents coverage
+3
Senior Application Reliability Engineer
Senior Application Reliability Engineer

Accenture in the Philippines • Taguig

On-site
PHP 1,200,000 - 2,200,000
Platform SRE Lead: Reliability, Automation & Scale
Platform SRE Lead: Reliability, Automation & Scale

Broadridge • Metro Manila

On-site
PHP 1,000,000 - 1,500,000
SRE: Automate, Scale & Stabilize Production Systems
SRE: Automate, Scale & Stabilize Production Systems

Philtech Inc. • Taguig

On-site
PHP 1,200,000 - 1,800,000
Site Reliability Engineer - Remote/Hybrid, High Impact
Site Reliability Engineer - Remote/Hybrid, High Impact

Manatal • Philippines

Hybrid
PHP 800,000 - 1,400,000
Healthcare coverage on day one
Dependents coverage
Paid time-off with cash conversion
+2
Staff SRE Engineer: AI-Driven Cloud Reliability Leader
Staff SRE Engineer: AI-Driven Cloud Reliability Leader

Stellar Cyber Inc. • España

On-site
PHP 1,200,000 - 1,800,000
Identity & Access SRE — Site Reliability Engineer
Identity & Access SRE — Site Reliability Engineer

Procter & Gamble Philippines • Manila

On-site
PHP 900,000 - 1,300,000
Performance bonus
Flexible work schedule / work fromHome
Health insurance & wellness programs
Site Reliability Engineer - Remote Philippines
Site Reliability Engineer - Remote Philippines

MicroSourcing • Manila

On-site
PHP 900,000 - 1,500,000
Healthcare coverage
Paid time-off with cash conversion
Group life insurance
+3