Hadoop Administrator & Data Engineer

Polus Solutions Pvt. Ltd.

Hinoba-an

On-site

PHP 1,800,000 - 3,000,000

Full time

27 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Polus Solutions Pvt. Ltd. is seeking a Senior Hadoop Administrator & Data Engineer to manage and optimize Hadoop clusters, including HDP/CDH environments and Spark workloads.

You’ll tune HDFS/YARN, supervise HiveServer2/Metastore, and administer Ambari, with strong Linux skills and scripting to support resilient data infrastructure. You will partner with data engineers to keep Airflow/AWS MWAA pipelines healthy, enforce capacity planning, and handle upgrades across on-prem and cloud-integrated

Qualifications

  • 4+ years administering production Hadoop clusters (HDP, CDH, or equivalent).
  • Deep, hands-on HDFS, YARN, and Hive administration — not just usage.
  • Strong Linux systems administration: SSH, systemd, and troubleshooting on RHEL/CentOS-family hosts.
  • Spark-on-YARN operational experience with memory tuning and driver issues.
  • Scripting proficiency in Bash and Python.
  • Experience with Ambari administration and in-place Hadoop version upgrades.
  • Airflow/AWS MWAA experience for data pipelines.
  • MySQL and MongoDB operations at scale.
  • Elasticsearch or Sphinx/Manticore search infrastructure familiarity.
  • Terraform and AWS (S3, EMR, IAM) expertise.
  • Experience decommissioning or migrating legacy Hadoop estates.

Responsibilities

  • Operate and tune HDFS, YARN, Hive, and Spark across clusters with mixed-version estates.
  • Manage gateway host capacity and prevent resource starvation.
  • Run Hive Metastore and HiveServer2 including JVM sizing and error diagnosis.
  • Administer clusters via Ambari: configs, alerts, and upgrades.
  • Plan capacity, quotas, and data retention with documented cleanup policies.
  • Collaborate with data engineering teams and support remote pipelines via SSH.
  • Maintain access control, SSH hygiene, and service accounts across nodes.
  • Build monitoring and alerting for host load, per-tenant usage, and metastore health.

Skills

HDFS administration
YARN administration
Hive administration
Linux administration
Bash scripting
Python scripting
Ambari administration
Apache Airflow
AWS MWAA
Terraform
SQL Databases (MySQL)
NoSQL (MongoDB)
Elasticsearch
Sphinx/Manticore
AWS Core services (S3, EMR, IAM)
Hadoop version migration

Tools

Hadoop (HDP/CDH)
Spark
Ambari
Airflow/AWS MWAA
MySQL
MongoDB
Elasticsearch
Terraform

Job description

This is a Senior Hadoop Administrator & Data Engineer role focused on managing and supporting Hadoop clusters and data infrastructure. The candidate should have strong hands-on experience with HDFS, YARN, Hive, Spark, Ambari, and Linux administration. Key responsibilities include performance tuning, capacity planning, monitoring, access control, troubleshooting, and managing version compatibility. They will also support Airflow/AWS MWAA pipelines and act as the escalation point for infrastructure-related failures.

Key Responsibilities
  • Operate and tune HDFS, YARN, Hive, and Spark across both clusters, including a mixed-version estate: one cluster runs an older HDP-era stack with Spark 2.x and a legacy Python driver environment; the other has recently been upgraded to a modern Hive release. Managing that version skew — and the compatibility traps it creates for job authors — is part of the role.
  • Own gateway host capacity. Monitor and set limits on concurrent worker pools, multiprocessing fan-outs, and Spark drivers so a single workload cannot starve the box. Push back on resource requests that the available evidence doesn’t actually support.
  • Run the Hive Metastore and HiveServer2, including JVM sizing and OOM exposure. Diagnose transient failures such as partition-add serialization errors and thrift session rejections.
  • Administer the clusters through Ambari: service configs, alerts, rolling restarts, and version upgrades.
  • Handle capacity planning, HDFS quotas, and data retention. Every dataset we generate needs a documented retention and cleanup story, and you’ll help enforce that on the storage side.
  • Partner with the data engineering team, which orchestrates pipelines with Apache Airflow on AWS MWAA. Most remote work reaches the clusters over SSH from the gateways, so you’ll be the escalation point when a pipeline task fails for infrastructure reasons rather than logic reasons.
  • Manage access control, SSH key hygiene, and service accounts across gateways and cluster nodes.
  • Build the monitoring and alerting we’re missing — host load, per-tenant resource consumption, metastore health, and early warning before a gateway becomes unresponsive.
Required Qualifications & Skills
  • Should work in US hours ( 6am Pacific Time to 5 Pacific Time )
  • 4+ years administering production Hadoop clusters (HDP, CDH, or equivalent).
  • Deep, hands-on HDFS, YARN, and Hive administration — not just usage. You should be comfortable reading NameNode and metastore logs and reasoning about JVM heap behavior.
  • Strong Linux systems administration: process and memory forensics, SSH, systemd, disk and network troubleshooting on RHEL/CentOS-family hosts.
  • Spark-on-YARN operational experience, including memory overhead tuning and diagnosing driver-side failures.
  • Solid scripting in Bash and Python.
  • Judgment about shared infrastructure. We want someone who will say “that benchmark measured database throughput, not gateway capacity — start lower and ramp” instead of maximizing a number.
  • Ambari administration and in-place Hadoop version upgrades.
  • Apache Airflow, particularly AWS MWAA.
  • MySQL and MongoDB operations at scale.
  • Elasticsearch or Sphinx/Manticore search infrastructure.
  • Terraform and AWS (S3, EMR, IAM).
  • Experience decommissioning or migrating legacy Hadoop estates onto newer platforms.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Hadoop & Data Infrastructure Engineer
Senior Hadoop & Data Infrastructure Engineer

Polus Solutions Pvt. Ltd. • Hinoba-an

On-site
PHP 1,800,000 - 3,000,000
Senior Data Platform Engineer
Senior Data Platform Engineer

Tookitaki • Manila

On-site
PHP 1,500,000 - 2,100,000
Data Platform Engineer
Data Platform Engineer

V2 Solutions • Hinoba-an

On-site
PHP 300,000 - 540,000
Pyspark Developer ( Mumbai)
Pyspark Developer ( Mumbai)

V2 Solutions • Hinoba-an

On-site
PHP 796,000 - 1,194,000
Data Engineer
Data Engineer

iGaming Centre • Manila

On-site
PHP 800,000 - 1,400,000
Data Analyst
Data Analyst

OpsWerks • Mandaluyong

On-site
PHP 800,000 - 1,200,000
Data Engineer
Data Engineer

Proselect Management Inc • Taguig

On-site
Remote Hadoop/Hive Administrator
Remote Hadoop/Hive Administrator

MCI • San Fernando

On-site
PHP 420,000 - 620,000
Lead Software Engineer – Data
Lead Software Engineer – Data

V2 Solutions • Hinoba-an

On-site
PHP 2,009,000 - 3,125,000
Platform Engineer - Data (SDE-3 / Staff)
Platform Engineer - Data (SDE-3 / Staff)

JobCubby • Hinoba-an

On-site
PHP 2,653,000 - 4,310,000