Data Engineer: Scalable Pipelines with Spark & Hive

Tata Consultancy Services

Irving (TX)

On-site

USD 125,000 - 140,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tata Consultancy Services in Irving, TX is seeking a Data Engineer to design, build, and optimize scalable data pipelines using Spark, PySpark, and Hive within Cloudera-like environments. Strong programming in Python, Spark SQL, and familiarity with Parquet/ORC formats are required, as is experience with Airflow or similar schedulers.

This role offers a competitive salary and opportunities to work on cutting-edge data platforms.

Qualifications

  • Proficient in Spark architecture, drivers, executors, and DAGs.
  • Strong Python and PySpark programming skills for complex transformations.
  • Strong HiveQL/ANSI SQL with partitioning and schema design.
  • Experience with Parquet/ORC/Avro storage formats.
  • Solid foundation in dimensional modeling and data lake concepts.

Responsibilities

  • Design, build, and maintain scalable ETL/ELT pipelines using PySpark and Spark SQL, Hive.
  • Optimize data warehouse layouts, partitioning, and indexing for performance.
  • Tune Spark jobs, monitor via Spark UI, and address memory issues and skew.
  • Ingest high-volume structured and unstructured data into the data ecosystem.
  • Automate data workflows with Airflow or platform schedulers for reliable delivery.
  • Collaborate with data scientists and analysts to translate business requirements into data solutions.

Skills

Spark Architecture
Python / PySpark
HiveQL / ANSI SQL
Parquet/ORC/Avro
Dimensional Modeling

Education

Bachelor's degree in Computer Science

Tools

Git
Jenkins
Ansible
AWS EMR
Databricks

Job description

Tata Consultancy Services in Irving, TX is seeking a Data Engineer to design, build, and optimize scalable data pipelines using Spark, PySpark, and Hive within Cloudera-like environments. Strong programming in Python, Spark SQL, and familiarity with Parquet/ORC formats are required, as is experience with Airflow or similar schedulers.

This role offers a competitive salary and opportunities to work on cutting-edge data platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer: Scalable Spark Pipelines & Cloud Infra
Data Engineer: Scalable Spark Pipelines & Cloud Infra

Tata Consultancy Services Limited • Irving (TX)

On-site
USD 70,000 - 80,000
Senior Data Engineer — Spark & Cloud Pipelines
Senior Data Engineer — Spark & Cloud Pipelines

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
Senior PySpark Data Engineer
Senior PySpark Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
Senior PySpark Engineer | Scalable ETL Pipelines & Cloud
Senior PySpark Engineer | Scalable ETL Pipelines & Cloud

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 150,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Parental Leaves
+4
Data Engineer: Spark, PySpark & Hive, Onsite Irving
Data Engineer: Spark, PySpark & Hive, Onsite Irving

Siri InfoSolutions Inc • Town of Texas (WI)

On-site
USD 110,000 - 160,000
Data Engineer
Data Engineer

Siri InfoSolutions Inc • Town of Texas (WI)

On-site
USD 110,000 - 160,000
Developer
Developer

Tata Consultancy Services Limited • Irving (TX)

On-site
USD 100,000 - 130,000
AWS PySpark Data Engineer: Scalable Data Pipelines
AWS PySpark Data Engineer: Scalable Data Pipelines

LTM • Irving (TX)

On-site
USD 120,000 - 180,000
Medical plan
Disability coverage
401(k) match
+3
Engineer
Engineer

Tata Consultancy Services Limited • Irving (TX)

On-site
USD 70,000 - 80,000
Senior PySpark Data Engineer - Cloud & ETL Pipelines
Senior PySpark Data Engineer - Cloud & ETL Pipelines

Tata Consultancy Services • Dallas (TX)

On-site
USD 100,000 - 105,000