Data Engineer: Scalable Pipelines with Spark & Hive

Tata Consultancy Services

Irving (TX)

On-site

USD 125,000 - 140,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Tata Consultancy Services in Irving, TX is seeking a Data Engineer to design, build, and optimize scalable data pipelines using Spark, PySpark, and Hive within Cloudera-like environments. Strong programming in Python, Spark SQL, and familiarity with Parquet/ORC formats are required, as is experience with Airflow or similar schedulers.

This role offers a competitive salary and opportunities to work on cutting-edge data platforms.

Qualifications

  • Proficient in Spark architecture, drivers, executors, and DAGs.
  • Strong Python and PySpark programming skills for complex transformations.
  • Strong HiveQL/ANSI SQL with partitioning and schema design.
  • Experience with Parquet/ORC/Avro storage formats.
  • Solid foundation in dimensional modeling and data lake concepts.

Responsibilities

  • Design, build, and maintain scalable ETL/ELT pipelines using PySpark and Spark SQL, Hive.
  • Optimize data warehouse layouts, partitioning, and indexing for performance.
  • Tune Spark jobs, monitor via Spark UI, and address memory issues and skew.
  • Ingest high-volume structured and unstructured data into the data ecosystem.
  • Automate data workflows with Airflow or platform schedulers for reliable delivery.
  • Collaborate with data scientists and analysts to translate business requirements into data solutions.

Skills

Spark Architecture
Python / PySpark
HiveQL / ANSI SQL
Parquet/ORC/Avro
Dimensional Modeling

Education

Bachelor's degree in Computer Science

Tools

Git
Jenkins
Ansible
AWS EMR
Databricks

Job description

Tata Consultancy Services in Irving, TX is seeking a Data Engineer to design, build, and optimize scalable data pipelines using Spark, PySpark, and Hive within Cloudera-like environments. Strong programming in Python, Spark SQL, and familiarity with Parquet/ORC formats are required, as is experience with Airflow or similar schedulers.

This role offers a competitive salary and opportunities to work on cutting-edge data platforms.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Cloud Data Engineer: Spark & Hive Data Pipelines
Cloud Data Engineer: Spark & Hive Data Pipelines

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 110,000
Spark Data Engineer (PySpark) - Cloud ETL & Pipelines
Spark Data Engineer (PySpark) - Cloud ETL & Pipelines

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 120,000
Discretionary Annual Incentive
Medical Coverage (Health, Dental & Vis
401K Plan
+1
Teradata Data Engineer: SQL, Spark & Data Pipelines
Teradata Data Engineer: SQL, Spark & Data Pipelines

Tata Consultancy Services • Charlotte (NC)

On-site
USD 100,000 - 105,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
401K Plan
+2
AWS PySpark Data Engineer: Scalable Data Pipelines
AWS PySpark Data Engineer: Scalable Data Pipelines

LTM • Irving (TX)

On-site
USD 120,000 - 180,000
Medical plan
Disability coverage
401(k) match
+3
Senior Data Engineer: PySpark, Airflow & GCP
Senior Data Engineer: PySpark, Airflow & GCP

Tata Consultancy Services • Sunnyvale (CA)

On-site
USD 110,000 - 150,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Family Support: Parental Leaves
+4
Data Engineer
Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 125,000 - 140,000
Senior PySpark Data Engineer - Hadoop & ETL
Senior PySpark Data Engineer - Hadoop & ETL

Tata Consultancy Services • Charlotte (NC)

On-site
USD 100,000 - 110,000
Discretionary annual incentive
Comprehensive medical coverage
Family leaves
+4
PySpark Data Engineer
PySpark Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 120,000
Discretionary Annual Incentive
Medical Coverage (Health, Dental & Vis
401K Plan
+1
Senior PySpark Data Engineer: ETL & Scalable Pipelines
Senior PySpark Data Engineer: ETL & Scalable Pipelines

Covetus • Irving (TX)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

The Value Maximizer • United States

On-site
USD 90,000 - 120,000