AWS PySpark Data Engineer: Scalable Data Pipelines

LTM

Irving (TX)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical plan
Disability coverage
401(k) match
Life insurance
Paid time off
Parental leave

Job summary

LTIMindtree in Irving, Texas seeks an AWS PySpark Developer to design, build, and optimize scalable data solutions on the AWS ecosystem. The role focuses on PySpark data processing, Iceberg/Hive data warehousing, and secure, cost-efficient pipelines.

You will deploy data solutions using S3, EMR, Glue, Athena, Lambda, and Redshift, while ensuring strong governance, performance, and collaboration with data scientists and engineers. This is a full-time, on-site position in the US.

Qualifications

  • AWS certification(s) required or in progress.
  • Hands-on with core AWS data services (S3, EC2, EMR, Glue, Athena, Lambda, Redshift, Kinesis).
  • Strong PySpark data processing and open table formats like Iceberg.
  • Experience with Hive-based data warehousing and data modeling.

Responsibilities

  • Design, implement, and deploy scalable data solutions on AWS.
  • Develop robust PySpark ETL/ELT pipelines for ingestion and loading.
  • Build and optimize data pipelines using Iceberg and Hive formats.
  • Design data warehousing schemas and optimize queries for analytics.
  • Configure secure cloud infrastructure, including VPCs and security groups.
  • Operate containerized workloads on EKS and manage Spark/Hive jobs.
  • Tune performance and manage cost of big data applications on AWS.
  • Implement data governance, security, and access controls in AWS.

Skills

AWS
PySpark
Data warehousing
Iceberg
Hive
ETL/ELT
Python
SQL
Boto3
Kubernetes

Tools

S3
EC2
EMR
Glue
Athena
Lambda
Redshift
Kinesis
Iceberg

Job description

LTIMindtree in Irving, Texas seeks an AWS PySpark Developer to design, build, and optimize scalable data solutions on the AWS ecosystem. The role focuses on PySpark data processing, Iceberg/Hive data warehousing, and secure, cost-efficient pipelines.

You will deploy data solutions using S3, EMR, Glue, Athena, Lambda, and Redshift, while ensuring strong governance, performance, and collaboration with data scientists and engineers. This is a full-time, on-site position in the US.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AWS Data Engineer: PySpark, Iceberg & Hive
Senior AWS Data Engineer: PySpark, Iceberg & Hive

LTM • Tampa (FL)

On-site
USD 140,000 - 190,000
Comprehensive Medical Plan
401(k) Plan with Company match
Paid Paternity and Maternity Leave
+2
Data Engineer: Scalable Spark Pipelines & Cloud Infra
Data Engineer: Scalable Spark Pipelines & Cloud Infra

Tata Consultancy Services Limited • Irving (TX)

On-site
USD 70,000 - 80,000
PySpark Data Engineer - Build Scalable Pipelines
PySpark Data Engineer - Build Scalable Pipelines

LTM • Tampa (FL)

On-site
USD 110,000 - 140,000
Medical plan
Disability coverage
401(k) with company match
+3
Senior Data Engineer — Spark & Cloud Pipelines
Senior Data Engineer — Spark & Cloud Pipelines

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
Data Engineer: Scalable Pipelines with Spark & Hive
Data Engineer: Scalable Pipelines with Spark & Hive

Tata Consultancy Services • Irving (TX)

On-site
USD 125,000 - 140,000
Data Engineer: Spark, PySpark & Hive, Onsite Irving
Data Engineer: Spark, PySpark & Hive, Onsite Irving

Siri InfoSolutions Inc • Town of Texas (WI)

On-site
USD 110,000 - 160,000
Senior PySpark Data Engineer
Senior PySpark Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
AWS PySpark Data Engineer - Python Pipelines
AWS PySpark Data Engineer - Python Pipelines

Polarits • Newark (NJ)

On-site
USD 140,000 - 190,000
Data Engineer - PySpark & AWS Cloud Pipelines
Data Engineer - PySpark & AWS Cloud Pipelines

SDLC Technologies • Charlotte (NC)

On-site
USD 90,000 - 150,000
Senior PySpark Data Engineer: ETL & Scalable Pipelines
Senior PySpark Data Engineer: ETL & Scalable Pipelines

Covetus • Irving (TX)

On-site
USD 100,000 - 130,000