Hadoop Administrator & Data Engineer

Qzigma Technologies

Thiruvananthapuram

Hybrid

INR 1,100,000 - 1,600,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Hybrid work model
Remote work after 6 months

Job summary

Qzigma Technologies in Thiruvananthapuram, Kerala is seeking an experienced Hadoop Administrator & Data Engineer to manage and optimize production Hadoop environments, working with the Data Engineering team.

This role covers HDFS, YARN, Hive and Spark clusters, Ambari, Airflow on AWS MWAA, Linux administration, and data pipelines. After an initial hybrid period, the position shifts to fully remote with a US shift.

Qualifications

  • 4+ years of hands-on experience administering production Hadoop clusters.
  • Strong experience with HDP, CDH, or equivalent Hadoop distributions.
  • Deep knowledge of HDFS, YARN, Hive, Spark.
  • Linux system administration, particularly RHEL/CentOS.
  • Experience with Spark-on-YARN memory tuning and driver failures.
  • Scripting in Bash and Python.
  • Ambari administration and Hadoop version upgrades.
  • Airflow, preferably AWS MWAA.
  • MySQL and MongoDB in production/large-scale environments.
  • Elasticsearch, Sphinx or Manticore preferred.

Responsibilities

  • Administer and optimize production HDFS, YARN, Hive, and Spark clusters.
  • Manage mixed Hadoop versions and troubleshoot compatibility.
  • Monitor cluster health, capacity, and resource utilization.
  • Administer Ambari services, alerts, restarts, and upgrades.
  • Support Airflow workflows and MWAA deployments.
  • Diagnose logs and perform root-cause analysis.

Skills

Hadoop Admin
HDFS
YARN
Hive
Spark
Linux Admin
Ambari
Airflow
AWS MWAA
MySQL
MongoDB
Elasticsearch
Scripting Bash
Python
Spark-on-YARN
Root-cause analysis

Tools

Apache Ambari
Apache Airflow
AWS MWAA
RHEL/CentOS

Job description

We are looking for an experienced Hadoop Administrator & Data Engineer to manage and optimize production Hadoop environments while working closely with the Data Engineering team. The role requires strong hands‑on expertise in HDFS, YARN, Hive, Spark, Linux administration, Ambari, and cloud‑based data pipelines.

The ideal candidate should be comfortable troubleshooting production infrastructure, managing mixed Hadoop versions, performing capacity planning, optimizing resource utilization, and supporting data pipelines running through Apache Airflow on AWS MWAA.


Key Responsibilities
Hadoop & Cluster Administration
  • Administer and optimize production HDFS, YARN, Hive, and Spark clusters.
  • Manage Hadoop environments with mixed software versions and troubleshoot compatibility issues.
  • Monitor cluster health, performance, capacity, and resource utilization.
  • Administer clusters using Apache Ambari, including service configurations, alerts, rolling restarts, and version upgrades.
  • Perform capacity planning, HDFS quota management, data retention, and storage cleanup.
  • Monitor and optimize NameNode, DataNode, ResourceManager, NodeManager, Hive Metastore, and HiveServer2 services.
Hive & Spark Administration
  • Manage Hive Metastore and HiveServer2, including JVM sizing and memory management.
  • Troubleshoot Hive failures such as partition‑related errors, serialization issues, and Thrift session failures.
  • Support Spark-on-YARN workloads and troubleshoot driver and executor failures.
  • Tune Spark memory overhead, resource allocation, concurrency, and workload performance.
  • Analyze NameNode, Hive Metastore, Spark, and other service logs to identify root causes.
Linux & Infrastructure Management
  • Perform hands‑on Linux administration on RHEL/CentOS-family systems.
  • Troubleshoot processes, memory utilization, disk capacity, network connectivity, and system performance.
  • Manage gateway hosts and ensure adequate capacity for concurrent workloads.
  • Monitor worker pools, multiprocessing workloads, Spark drivers, and gateway resource consumption.
  • Establish appropriate resource limits to prevent individual workloads from impacting shared infrastructure.
  • Manage SSH access, SSH key hygiene, service accounts, and permissions across gateways and cluster nodes.
Monitoring & Performance
  • Develop and enhance monitoring and alerting for:
    • Host load and system health
    • Per‑tenant resource consumption
    • HDFS capacity and utilization
    • Hive Metastore health
    • Gateway capacity
    • Cluster services and availability
  • Identify early warning indicators and proactively resolve infrastructure issues.
  • Perform capacity planning and recommend resource allocation based on actual workload evidence and performance data.
Data Engineering & Cloud Support
  • Work closely with Data Engineering teams to support production data pipelines.
  • Provide infrastructure‑level troubleshooting for Apache Airflow workflows running on AWS MWAA.
  • Diagnose whether pipeline failures are related to infrastructure, cluster configuration, resource constraints, or application logic.
  • Support data engineering teams through SSH‑based access from gateway hosts.
  • Collaborate with developers and data engineers to ensure reliable and scalable data processing.
Database & Search Infrastructure
  • Support operations and troubleshooting for MySQL and MongoDB environments at scale.
  • Work with Elasticsearch or Sphinx/Manticore search infrastructure.
  • Monitor performance, troubleshoot issues, and coordinate infrastructure improvements as required.
Required Skills & Experience
  • 4+ years of hands‑on experience administering production Hadoop clusters.
  • Strong experience with HDP, CDH, or equivalent Hadoop distributions.
  • Deep knowledge of:
    • HDFS
    • YARN
    • Hive
    • Spark
  • Strong experience in Linux system administration, particularly RHEL/CentOS.
  • Experience with process and memory troubleshooting, SSH, systemd, disk, and network troubleshooting.
  • Hands‑on experience with Spark-on-YARN, including memory overhead tuning and driver‑side failure troubleshooting.
  • Strong scripting skills in Bash and Python.
  • Hands‑on experience with Apache Ambari administration and Hadoop version upgrades.
  • Experience with Apache Airflow, preferably AWS MWAA.
  • Experience supporting MySQL and MongoDB in production/large‑scale environments.
  • Experience with Elasticsearch, Sphinx, or Manticore is preferred.
  • Strong analytical and troubleshooting skills with the ability to perform root‑cause analysis.
  • Good understanding of shared infrastructure, resource management, capacity planning, and workload optimization.
Preferred Candidate Profile
  • Strong production troubleshooting and incident‑management experience.
  • Ability to analyze logs, JVM behavior, memory utilization, and system performance.
  • Comfortable working with large‑scale distributed data platforms.
  • Strong judgment regarding infrastructure capacity and resource requests.
  • Proactive approach toward monitoring, alerting, optimization, and preventive maintenance.
  • Good communication and collaboration skills with Data Engineering and technical teams.
  • Ability to work independently in a remote/hybrid environment.
Work Schedule & Location
  • Location: Thiruvananthapuram, Kerala
  • First 6 Months: Hybrid work model
  • After 6 Months: Fully Remote
  • Shift: US Shift
  • Candidate must be comfortable working during 6:00 AM5:00 PM Pacific Time business hours.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hadoop Administrator & Data Engineer
Hadoop Administrator & Data Engineer

Polus Solutions • Thiruvananthapuram

Hybrid
INR 1,200,000 - 1,800,000
Senior Big Data Platform Administrator
Senior Big Data Platform Administrator

Larsen & Toubro • Chennai District

On-site
INR 2,500,000 - 3,800,000
Hadoop Administrator
Hadoop Administrator

Infosys • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Hadoop Administrator(Admin)
Hadoop Administrator(Admin)

Infosys • Hyderabad, Pune District, Bengaluru

On-site
INR 600,000 - 900,000
Bigdata Developer
Bigdata Developer

Unison Group • Hyderabad

On-site
INR 800,000 - 1,500,000
Hadoop Admin
Hadoop Admin

Capco • Bengaluru

On-site
INR 1,500,000 - 2,700,000
Big Data Developer
Big Data Developer

Viraaj HR Solutions Private Limited • Maharashtra

On-site
INR 1,000,000 - 1,500,000
Collaborative engineering culture
Competitive compensation
Training in cloud and Big Data technologies
Data Platform Engineer
Data Platform Engineer

Yotta • Delhi

On-site
INR 600,000 - 1,000,000
Support Analyst
Support Analyst

Ex • Pune District

On-site
INR 600,000 - 1,200,000
Hadoop Administrator
Hadoop Administrator

Hackajob • Hyderabad, Pune District, Bengaluru

Hybrid
INR 3,500,000 - 4,500,000