Senior PySpark Data Engineer

Tata Consultancy Services

Irving (TX)

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tata Consultancy Services in Irving, TX is seeking a Data Engineer to design, build, and optimize scalable data pipelines using Spark, PySpark, and Hive in a cloud-first environment.

The role emphasizes data reliability, speed, and analytics readiness, collaborating with data scientists and analysts to translate business requirements into robust solutions. Experience with AWS/Azure/GCP data services is valued.

Qualifications

  • Big data frameworks expertise with Spark architecture, drivers, executors, and DAGs.
  • Advanced programming: Python and PySpark API for data transformations.
  • Querying & schema management: HiveQL and ANSI SQL, with partitioning and schema definition.
  • Optimized storage formats: Parquet, ORC, Avro.
  • Cloud ecosystem development: AWS EMR, Azure Databricks, cloud data services.
  • Data warehousing fundamentals: Star/Snowflake schemas and data lakes concepts.

Responsibilities

  • Data Pipeline Development & Maintenance: Design, build, and maintain scalable ETL/ELT pipelines using PySpark and Spark SQL.
  • Cloud Data Infrastructure Management: Deploy and scale data infrastructure on AWS, Azure, or GCP.
  • Data Warehousing & Storage Optimization: Manage data layout, partitioning, indexing in Hive and cloud data lakes.
  • Performance Tuning: Identify and resolve Spark job bottlenecks and memory issues.
  • Diverse Data Integration: Ingest high-volume datasets from relational and unstructured sources.
  • Automated Workflow Orchestration: Implement data workflows with Airflow or schedulers.
  • Strategic Collaboration: Work with data scientists and analysts to translate requirements.

Skills

Apache Spark
PySpark API
Python
Spark SQL
HiveQL
ANSI SQL
Parquet/ORC/Avro
AWS EMR
Azure Databricks
Airflow
Data Modeling

Education

Bachelor's degree in Computer Science

Tools

AWS EMR
Azure Databricks
GCP (BigQuery)
Airflow

Job description

Tata Consultancy Services in Irving, TX is seeking a Data Engineer to design, build, and optimize scalable data pipelines using Spark, PySpark, and Hive in a cloud-first environment.

The role emphasizes data reliability, speed, and analytics readiness, collaborating with data scientists and analysts to translate business requirements into robust solutions. Experience with AWS/Azure/GCP data services is valued.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer — Spark & Cloud Pipelines
Senior Data Engineer — Spark & Cloud Pipelines

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
Data Engineer: Scalable Spark Pipelines & Cloud Infra
Data Engineer: Scalable Spark Pipelines & Cloud Infra

Tata Consultancy Services Limited • Irving (TX)

On-site
USD 70,000 - 80,000
Senior PySpark Engineer | Scalable ETL Pipelines & Cloud
Senior PySpark Engineer | Scalable ETL Pipelines & Cloud

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 150,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Parental Leaves
+4
Data Engineer: Scalable Pipelines with Spark & Hive
Data Engineer: Scalable Pipelines with Spark & Hive

Tata Consultancy Services • Irving (TX)

On-site
USD 125,000 - 140,000
Senior Data Engineer - Azure SQL, Spark & ETL Expert
Senior Data Engineer - Azure SQL, Spark & ETL Expert

Tata Consultancy Services • Irving (TX)

On-site
USD 120,000 - 130,000
Discretionary Annual Incentive
Medical Coverage: Medical & Health, +
Family Leaves
+4
Senior PySpark Data Engineer - Cloud & ETL Pipelines
Senior PySpark Data Engineer - Cloud & ETL Pipelines

Tata Consultancy Services • Dallas (TX)

On-site
USD 100,000 - 105,000
Data Engineer: Spark, PySpark & Hive, Onsite Irving
Data Engineer: Spark, PySpark & Hive, Onsite Irving

Siri InfoSolutions Inc • Town of Texas (WI)

On-site
USD 110,000 - 160,000
Senior Data Engineer: Azure Data Pipelines & PySpark Expert
Senior Data Engineer: Azure Data Pipelines & PySpark Expert

Tata Consultancy Services • Chicago (IL)

On-site
USD 100,000 - 120,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Maternal & Parental Leaves
+2
AWS PySpark Data Engineer: Scalable Data Pipelines
AWS PySpark Data Engineer: Scalable Data Pipelines

LTM • Irving (TX)

On-site
USD 120,000 - 180,000
Medical plan
Disability coverage
401(k) match
+3
Senior PySpark Data Engineer: ETL & Scalable Pipelines
Senior PySpark Data Engineer: ETL & Scalable Pipelines

Covetus • Irving (TX)

On-site
USD 100,000 - 130,000