Hadoop Administrator & Data Engineer

Polus Solutions

Thiruvananthapuram

Hybrid

INR 1,200,000 - 1,800,000

Full time

12 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Polus Solutions is seeking an experienced Hadoop administrator to operate and tune HDFS, YARN, Hive, and Spark across mixed-version clusters. You will manage capacity, metastore services, and gateway resources while ensuring stable performance and data retention policies.

Responsibilities include administering Ambari-led clusters, coordinating with the data engineering team using AWS MWAA, and performing in-place upgrades. Strong Linux, Bash, and Python scripting are essential.

Qualifications

  • 4+ years administering production Hadoop clusters (HDP, CDH, or equivalent).
  • Comfort reading NameNode and metastore logs and reasoning about JVM heap behavior.
  • Strong Linux systems administration: SSH, systemd, disk and network troubleshooting on RHEL/CentOS-family hosts.

Responsibilities

  • Operate and tune HDFS, YARN, Hive, and Spark across clusters with mixed-version estate and legacy Python drivers.
  • Own gateway host capacity and set limits to prevent resource starvation.
  • Run Hive Metastore and HiveServer2; diagnose partition-add serialization errors and thrift session rejections.

Skills

HDFS
YARN
Hive
Spark
Ambari
Linux admin
Bash
Python
Terraform
AWS MWAA
MySQL
MongoDB
Elasticsearch
SSH
IAM
JVM tuning
Logging/diagnostics

Tools

Ambari
Airflow
AWS S3
EMR

Job description

Role & responsibilities
  • Operate and tune HDFS, YARN, Hive, and Spark across both clusters, including a mixed-version estate: one cluster runs an older HDP-era stack with Spark 2.x and a legacy Python driver environment; the other has recently been upgraded to a modern Hive release. Managing that version skew and the compatibility traps it creates for job authors is part of the role.
  • Own gateway host capacity. Monitor and set limits on concurrent worker pools, multiprocessing fan-outs, and Spark drivers so a single workload cannot starve the box. Push back on resource requests that the available evidence doesn't actually support.
  • Run the Hive Metastore and HiveServer2, including JVM sizing and OOM exposure. Diagnose transient failures such as partition-add serialization errors and thrift session rejections.
  • Administer the clusters through Ambari: service configs, alerts, rolling restarts, and version upgrades.
  • Handle capacity planning, HDFS quotas, and data retention. Every dataset we generate needs a documented retention and cleanup story, and you'll help enforce that on the storage side.
  • Partner with the data engineering team, which orchestrates pipelines with Apache Airflow on AWS MWAA. Most remote work reaches the clusters over SSH from the gateways, so you'll be the escalation point when a pipeline task fails for infrastructure reasons rather than logic reasons.
  • Manage access control, SSH key hygiene, and service accounts across gateways and cluster nodes.
  • Build the monitoring and alerting we're missing - host load, per-tenant resource consumption, metastore health, and early warning before a gateway becomes unresponsive.
Preferred candidate profile
  • Should work in US hours ( 6am Pacific Time to 5 Pacific Time )
  • 4+ years administering production Hadoop clusters (HDP, CDH, or equivalent).
  • Deep, hands-on HDFS, YARN, and Hive administration - not just usage. You should be comfortable reading NameNode and metastore logs and reasoning about JVM heap behavior.
  • Strong Linux systems administration: process and memory forensics, SSH, systemd, disk and network troubleshooting on RHEL/CentOS-family hosts.
  • Spark-on-YARN operational experience, including memory overhead tuning and diagnosing driver-side failures.
  • Solid scripting in Bash and Python.
  • Judgment about shared infrastructure. We want someone who will say 'that benchmark measured database throughput, not gateway capacity - start lower and ramp' instead of maximizing a number.
  • Ambari administration and in-place Hadoop version upgrades.
  • Apache Airflow, particularly AWS MWAA.
  • MySQL and MongoDB operations at scale.
  • Elasticsearch or Sphinx/Manticore search infrastructure.
  • Terraform and AWS (S3, EMR, IAM).
  • Experience decommissioning or migrating legacy Hadoop estates onto newer platforms.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hadoop Administrator
Hadoop Administrator

Infosys • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Big Data Administrator
Big Data Administrator

Capgemini • Bengaluru, Hyderabad

Hybrid
INR 900,000 - 1,300,000
Hadoop Administrator(Admin)
Hadoop Administrator(Admin)

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 600,000 - 900,000
Hadoop Administrator & Data Engineer
Hadoop Administrator & Data Engineer

Qzigma Technologies • Thiruvananthapuram

Hybrid
INR 1,100,000 - 1,600,000
Hybrid work model
Remote work after 6 months
Hadoop Administrator
Hadoop Administrator

Paytm • Dadri

On-site
INR 800,000 - 1,200,000
Hadoop Administrator
Hadoop Administrator

Hackajob • Hyderabad, Pune District, Bengaluru

Hybrid
INR 3,500,000 - 4,500,000
Senior Big Data Platform Administrator
Senior Big Data Platform Administrator

Larsen & Toubro • Chennai District

On-site
INR 2,500,000 - 3,800,000
Senior Hadoop Administrator
Senior Hadoop Administrator

Persistent • Pune District

On-site
INR 1,200,000 - 2,400,000
Competitive salary and benefits
Growth opportunities with education/cs
Cutting-edge technologies
+3
Hadoop Developer
Hadoop Developer

Cloudxtreme • Pune District

On-site
INR 4,000,000 - 6,000,000
Cloudera Admin
Cloudera Admin

Unison Group • Mumbai

On-site
INR 600,000 - 1,200,000