PySpark Consultant

VDart Inc

Irving (TX)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

VDart Inc. in Irving, TX is seeking an experienced Big Data Developer to design and implement PySpark-based data processing pipelines. You will ensure data quality, troubleshoot PySpark workflows, and integrate with ingestion and data-lens frameworks while maintaining data security standards.

You should have 4+ years in big data with Hadoop, Hive and Spark, strong Python/SQL skills, and a proven ability to lead, communicate, and document data lineage across complex data flows.

Qualifications

  • 4+ years of experience in big data development, Hadoop, Hive & Spark.
  • Strong Python and SQL knowledge.
  • Experience with PySpark and Spark frameworks.
  • Knowledge of data security and CI/CD practices.

Responsibilities

  • Design and implement ETL pipelines and data transformations.
  • Troubleshoot PySpark applications and workflows.
  • Ensure data quality and integrity across processing workflows.
  • Document data lineage and data flow.
  • Understand sources and dependencies of converted PySpark code.
  • Integrate PySpark code with Ingestion Framework and DataLens where applicable.

Skills

Big data development
Python
SQL
PySpark
Spark framework
Hadoop
Hive
Kafka
Data lineage
CI/CD / DevOps
Leadership
Communication

Tools

PySpark
Spark
Hadoop
Hive
Kafka

Job description

Experience with big data processing and distributed computing systems like Spark.

Implement ETL pipelines and data transformation processes.

Ensure data quality and integrity in all data processing workflows.

Troubleshoot and resolve issues related to PySpark applications and workflows.

Understand source, dependencies and data flow from converted PySpark code.

Strong programming skills in Python and SQL.

Experience with big data technologies like Hadoop, Hive, and Kafka.

Understanding of data warehousing concepts and relational databases like SQL.

Demonstrate and document code lineage.

Integrate PySpark code with frameworks such as Ingestion Framework, DataLens, etc.,

Ensure compliance with data security, privacy regulations, and organizational standards.

Knowledge of CI/CD pipelines and DevOps practices.

Strong problem-solving and analytical skills.

Excellent communication and leadership abilities. Qualifications:

4+ years of experience in big data development, Hadoop , Hive & Spark framework.

Good to have experience in SAS.

Strong Python, PySpark Development and SQL knowledge.

Certification in big data or cloud technologies is preferred.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

PySpark Consultant
PySpark Consultant

VDart • Irving (TX)

On-site
USD 120,000 - 150,000
PySpark Consultant Contract C2C jobs Urgent Need
PySpark Consultant Contract C2C jobs Urgent Need

Tech Mirrors • Irving (TX)

On-site
USD 96,000 - 152,000
Pyspark Developer
Pyspark Developer

Tieto • Irving (TX)

On-site
USD 140,000 - 190,000
Senior PySpark Data Engineer
Senior PySpark Data Engineer

Covetus • Irving (TX)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

The Value Maximizer • United States

On-site
USD 90,000 - 120,000
PySpark Data Engineer
PySpark Data Engineer

Tata Consultancy Services • Irving (TX)

On-site
USD 90,000 - 120,000
Discretionary Annual Incentive
Medical Coverage (Health, Dental & Vis
401K Plan
+1
Hadoop developer
Hadoop developer

REALIGN LLC • Charlotte (NC)

On-site
USD 120,000 - 180,000
Pyspark developer (Data engineer)
Pyspark developer (Data engineer)

MTK Technologies • Irving (TX)

On-site
USD 90,000 - 120,000
Data Engineer ETL
Data Engineer ETL

Compunnel, Inc. • Durham (NC)

On-site
USD 100,000 - 130,000
Pyspark Architect
Pyspark Architect

Avance Consulting • Charlotte (NC)

On-site
USD 100,000 - 130,000