PySpark / Java Developer

Veriipro

Whitpain Township (PA)

On-site

USD 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veriipro in Whitpain Township, Pennsylvania, is looking for an experienced data engineer to design and maintain ETL pipelines for structured and unstructured data. You will work with technologies like PySpark, SQL Server, and Hadoop, ensuring data pipeline reliability and performance.

The ideal candidate should have over 5 years of experience in database development and a strong understanding of data processing applications. Join us to leverage your expertise in a dynamic environment.

Qualifications

  • 5+ years of experience in Microsoft SQL Server and relational database development.
  • Strong understanding of ETL concepts and large-scale data processing.
  • Proven experience in performance tuning of SQL queries.

Responsibilities

  • Design, develop, and maintain scalable ETL pipelines.
  • Build and optimize data processing applications using PySpark and Java.
  • Analyze and resolve performance bottlenecks in SQL procedures.

Skills

Microsoft SQL Server
ETL concepts
SQL performance tuning
Hadoop
PySpark
Hive
Impala
Python

Tools

Cloudera Hadoop Ecosystem
Kafka
Hue
Oozie
YARN
Sqoop

Job description

Roles and Responsibilities
  • Design, develop, and maintain scalable ETL pipelines for large-scale structured and unstructured data.
  • Build and optimize data processing applications using PySpark and Java.
  • Work extensively with relational databases and big data platforms for data extraction, transformation, and loading.
  • Analyze and resolve performance bottlenecks in high-volume SQL procedures and big data processing jobs.
  • Develop efficient data movement and transformation workflows across distributed systems.
  • Collaborate with cross-functional teams to understand end-to-end data flow and business requirements.
  • Support production systems, troubleshoot issues, and ensure data pipeline reliability.
Required Skills & Experience
  • 5+ years of experience in Microsoft SQL Server and relational database development for data extraction applications.
  • Strong understanding of ETL concepts, database technologies, and large-scale data processing.
  • Proven experience in performance tuning of SQL queries and understanding of different indexing strategies.
  • 2+ years of experience working with big data technologies including:
    • Hadoop
    • Spark / PySpark
    • Hive
    • Impala
    • Python
  • 2+ years of hands‑on experience with the Cloudera Hadoop Ecosystem, including:
    • HDFS
    • Hive
    • Impala
    • Spark
    • Kafka
    • Hue
    • Oozie
    • YARN
    • Sqoop
  • Experience in processing large volumes of structured and unstructured data using Spark.
  • Strong understanding of end-to-end (E2E) data pipeline architecture and application workflows.
Preferred Skills
  • Domain experience in healthcare claims data or healthcare analytics.
  • Experience with distributed data processing and optimization in production environments.
  • Strong troubleshooting and analytical skills in complex data ecosystems.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

PySpark Developer
PySpark Developer

Inizio Partners Corp • Hartford (CT)

On-site
USD 90,000 - 120,000
Data Engineer/Python Developer
Data Engineer/Python Developer

TechDigital Group • Minnesota

On-site
USD 80,000 - 120,000
Pyspark Developer
Pyspark Developer

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
Software developer
Software developer

Compunnel, Inc. • Atlanta (GA)

On-site
USD 80,000 - 110,000
Java Spark Engineer
Java Spark Engineer

Veriipro • Berkeley Heights (NJ)

On-site
USD 140,000 - 190,000
Data Engineer
Data Engineer

Jobtailor • Kentucky

On-site
USD 110,000 - 140,000
Developer
Developer

Tata Consultancy Services • Irving (TX)

On-site
USD 100,000 - 130,000
Senior PySpark Data Engineer
Senior PySpark Data Engineer

Covetus • Irving (TX)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

SDLC Technologies • Charlotte (NC)

On-site
USD 90,000 - 150,000
Data Engineer ETL
Data Engineer ETL

Compunnel, Inc. • Durham (NC)

On-site
USD 100,000 - 130,000