This is a Senior Hadoop Administrator & Data Engineer role focused on managing and supporting Hadoop clusters and data infrastructure. The candidate should have strong hands-on experience with HDFS, YARN, Hive, Spark, Ambari, and Linux administration. Key responsibilities include performance tuning, capacity planning, monitoring, access control, troubleshooting, and managing version compatibility. They will also support Airflow/AWS MWAA pipelines and act as the escalation point for infrastructure-related failures.
Key Responsibilities
- Operate and tune HDFS, YARN, Hive, and Spark across both clusters, including a mixed-version estate: one cluster runs an older HDP-era stack with Spark 2.x and a legacy Python driver environment; the other has recently been upgraded to a modern Hive release. Managing that version skew — and the compatibility traps it creates for job authors — is part of the role.
- Own gateway host capacity. Monitor and set limits on concurrent worker pools, multiprocessing fan-outs, and Spark drivers so a single workload cannot starve the box. Push back on resource requests that the available evidence doesn’t actually support.
- Run the Hive Metastore and HiveServer2, including JVM sizing and OOM exposure. Diagnose transient failures such as partition-add serialization errors and thrift session rejections.
- Administer the clusters through Ambari: service configs, alerts, rolling restarts, and version upgrades.
- Handle capacity planning, HDFS quotas, and data retention. Every dataset we generate needs a documented retention and cleanup story, and you’ll help enforce that on the storage side.
- Partner with the data engineering team, which orchestrates pipelines with Apache Airflow on AWS MWAA. Most remote work reaches the clusters over SSH from the gateways, so you’ll be the escalation point when a pipeline task fails for infrastructure reasons rather than logic reasons.
- Manage access control, SSH key hygiene, and service accounts across gateways and cluster nodes.
- Build the monitoring and alerting we’re missing — host load, per-tenant resource consumption, metastore health, and early warning before a gateway becomes unresponsive.
Required Qualifications & Skills
- Should work in US hours ( 6am Pacific Time to 5 Pacific Time )
- 4+ years administering production Hadoop clusters (HDP, CDH, or equivalent).
- Deep, hands-on HDFS, YARN, and Hive administration — not just usage. You should be comfortable reading NameNode and metastore logs and reasoning about JVM heap behavior.
- Strong Linux systems administration: process and memory forensics, SSH, systemd, disk and network troubleshooting on RHEL/CentOS-family hosts.
- Spark-on-YARN operational experience, including memory overhead tuning and diagnosing driver-side failures.
- Solid scripting in Bash and Python.
- Judgment about shared infrastructure. We want someone who will say “that benchmark measured database throughput, not gateway capacity — start lower and ramp” instead of maximizing a number.
- Ambari administration and in-place Hadoop version upgrades.
- Apache Airflow, particularly AWS MWAA.
- MySQL and MongoDB operations at scale.
- Elasticsearch or Sphinx/Manticore search infrastructure.
- Terraform and AWS (S3, EMR, IAM).
- Experience decommissioning or migrating legacy Hadoop estates onto newer platforms.