Data Engineer: Spark, PySpark & Hive on Cloudera

Prism IT Global

Irving (TX)

On-site

USD 120,000 - 150,000

Full time

22 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Prism IT Global is seeking a Data Engineer to design, build, and optimize scalable data pipelines using Apache Spark, PySpark, and Hive within Cloudera/AWS environments.

You will develop ETL/ELT processes, manage data warehousing structures, and ensure high performance through tuning, partitioning, and efficient storage formats. Collaboration with data scientists and analytics teams is essential.

Qualifications

  • Proficiency with Spark architecture, drivers, executors and DAGs.
  • Strong Python programming with PySpark for data transformations.
  • Experience with HiveQL/ANSI SQL and partitioning strategies.
  • Knowledge of optimized big data formats (Parquet/ORC/Avro).
  • Solid understanding of dimensional modeling and data warehouse concepts.

Responsibilities

  • Design, build and maintain scalable ETL/ELT pipelines using PySpark and Spark SQL/Hive.
  • Optimize data layout, partitioning and indexing for performance.
  • Tune Spark jobs and manage memory to reduce skew and bottlenecks.
  • Ingest high-volume structured and unstructured data into the data ecosystem.
  • Implement automated workflows with Airflow or similar schedulers.
  • Collaborate with data scientists and analysts to translate requirements into data solutions.

Skills

Apache Spark
PySpark
HiveQL
Python
Data modeling
ETL

Tools

Parquet/ORC/Avro
Spark SQL
Hive
Airflow
AWS EMR
Databricks

Job description

Prism IT Global is seeking a Data Engineer to design, build, and optimize scalable data pipelines using Apache Spark, PySpark, and Hive within Cloudera/AWS environments.

You will develop ETL/ELT processes, manage data warehousing structures, and ensure high performance through tuning, partitioning, and efficient storage formats. Collaboration with data scientists and analytics teams is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer (Cloudera Platform)
Data Engineer (Cloudera Platform)

Prism IT Global • Irving (TX)

On-site
USD 120,000 - 150,000
Data Engineer
Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 125,000 - 140,000
Data Engineer: Scalable Pipelines with Spark & Hive
Data Engineer: Scalable Pipelines with Spark & Hive

Tata Consultancy Services • Irving (TX)

On-site
USD 125,000 - 140,000
Big Data Engineer: Spark, ETL & Data Quality
Big Data Engineer: Spark, ETL & Data Quality

eNcloud Services LLC • Katy (TX)

On-site
USD 100,000 - 170,000
Data Engineer
Data Engineer

The Value Maximizer • United States

On-site
USD 90,000 - 120,000
PySpark Data Engineer
PySpark Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 120,000
Discretionary Annual Incentive
Medical Coverage (Health, Dental & Vis
401K Plan
+1
Data Engineer ETL
Data Engineer ETL

Compunnel, Inc. • Durham (NC)

On-site
USD 100,000 - 130,000
Senior PySpark Engineer | Scalable ETL Pipelines & Cloud
Senior PySpark Engineer | Scalable ETL Pipelines & Cloud

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 150,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Parental Leaves
+4
Data Engineer: Snowflake, Spark & ETL for Analytics
Data Engineer: Snowflake, Spark & ETL for Analytics

Praise Tech Solutions • Northern (KY)

Hybrid
USD 120,000 - 160,000
Data Engineer: Cloud Data Pipelines & Lakehouse
Data Engineer: Cloud Data Pipelines & Lakehouse

Creative Information Technology India • Falls Church (VA)

On-site
USD 120,000 - 160,000