PySpark Data Engineer - Build Scalable Pipelines

LTM

Tampa (FL)

On-site

USD 110,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical plan
Disability coverage
401(k) with company match
Life insurance
Parental leave
Paid vacation and holidays

Job summary

LTIMindtree in Tampa, FL is seeking a PySpark Developer to design, build and optimize large-scale data pipelines using PySpark and Spark SQL. You will work with Azure Databricks, ADF, Synapse and Hadoop ecosystems, integrate with DataOps practices, and collaborate with data engineers, analysts and DevOps in an Agile environment.

Strong SQL, Linux, ETL fundamentals and Python scripting are required, plus a banking analytics background is a plus.

Qualifications

  • Hands-on PySpark and Apache Spark experience.
  • Strong SQL and relational database knowledge.
  • Experience in Linux/Unix environments.
  • ETL and large-scale data processing.
  • Exposure to Azure Databricks, ADF, Synapse or similar.
  • Hadoop ecosystem (Hive, HDFS).
  • Airflow orchestration knowledge.
  • Python scripting beyond Spark usage.
  • Agile and DevOps delivery practices.
  • Banking/BFSI analytics exposure is a plus.
  • Strong problem-solving and ownership mindset.
  • Cross-functional collaboration and communication skills.

Responsibilities

  • Design, develop and optimize data pipelines using PySpark for large-scale processing.
  • Transform, validate and aggregate data with Spark SQL and DataFrames.
  • Write and optimize SQL queries for data extraction and reporting.
  • Perform data quality checks and profiling.
  • Troubleshoot production pipelines and perform root-cause analysis.
  • Collaborate with data engineers, analysts, QA and DevOps in Agile teams.
  • Enhance Spark job performance for scalability and efficiency.
  • Maintain documentation, runbooks and technical workflows.
  • Adhere to enterprise coding, deployment and change-management standards.
  • Self-direct learning of new big-data technologies and cloud platforms.

Skills

PySpark
SQL
Linux/Unix
ETL
Airflow
Python scripting
Agile/DevOps
Spark SQL
Data processing
Banking analytics (BFSI)

Tools

Azure Databricks
ADF
Synapse
Hadoop/Hive/HDFS
Airflow

Job description

LTIMindtree in Tampa, FL is seeking a PySpark Developer to design, build and optimize large-scale data pipelines using PySpark and Spark SQL. You will work with Azure Databricks, ADF, Synapse and Hadoop ecosystems, integrate with DataOps practices, and collaborate with data engineers, analysts and DevOps in an Agile environment.

Strong SQL, Linux, ETL fundamentals and Python scripting are required, plus a banking analytics background is a plus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Architect: AI-Enablement & Scalable Data Pipelines
Data Architect: AI-Enablement & Scalable Data Pipelines

LTM • Tampa (FL)

On-site
USD 120,000 - 180,000
Comprehensive Medical Plan
Disability Coverage
401(k) Plan with Company match
+3
AWS PySpark Data Engineer: Scalable Data Pipelines
AWS PySpark Data Engineer: Scalable Data Pipelines

LTM • Irving (TX)

On-site
USD 120,000 - 180,000
Medical plan
Disability coverage
401(k) match
+3
Data Engineer: Scalable Spark Pipelines & Cloud Infra
Data Engineer: Scalable Spark Pipelines & Cloud Infra

Tata Consultancy Services Limited • Irving (TX)

On-site
USD 70,000 - 80,000
Azure Databricks Lead: Pipelines, Delta Lake & Spark Expert
Azure Databricks Lead: Pipelines, Delta Lake & Spark Expert

Covetus • New York (NY)

On-site
USD 130,000 - 175,000
Senior PySpark Engineer | Scalable ETL Pipelines & Cloud
Senior PySpark Engineer | Scalable ETL Pipelines & Cloud

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 150,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Parental Leaves
+4
Azure Databricks Engineer - PySpark & ETL Specialist
Azure Databricks Engineer - PySpark & ETL Specialist

LTM • Branchville (NJ)

On-site
USD 120,000 - 160,000
Comprehensive Medical Plan Covering .M
Short Term and Long-Term Disability .?
401(k) Plan with Company match
+2
Specialist - Data Engineering
Specialist - Data Engineering

LTM • Tampa (FL)

On-site
USD 110,000 - 140,000
Medical plan
Disability coverage
401(k) with company match
+3
Senior Data Engineer — Spark & Cloud Pipelines
Senior Data Engineer — Spark & Cloud Pipelines

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
Senior Data Engineer: Azure Data Pipelines & PySpark Expert
Senior Data Engineer: Azure Data Pipelines & PySpark Expert

Tata Consultancy Services • Chicago (IL)

On-site
USD 100,000 - 120,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Maternal & Parental Leaves
+2
Data Engineer: Scalable Pipelines with Spark & Hive
Data Engineer: Scalable Pipelines with Spark & Hive

Tata Consultancy Services • Irving (TX)

On-site
USD 125,000 - 140,000