Lead Databricks Engineer ( F2F interview | Parsippany, NJ | only W2 )

COOLSOFT

Parsippany-Troy Hills (NJ)

On-site

USD 140,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

COOLSOFT in Parsippany-Troy Hills, NJ seeks a Lead Data Engineer to design and maintain scalable data pipelines for payroll and financial datasets, leveraging Databricks, PySpark, Python, and SQL.

You will collaborate with data scientists and stakeholders to deliver governance, lakehouse architecture, and analytics solutions, with a 70% data engineering and 30% analytics mix.

Qualifications

  • Bachelor's or Master's degree in CS, Data Eng, or related field.
  • 5+ years of experience in Data Engineering or Data Platform Development.
  • Strong hands-on experience with Databricks, PySpark, Python, and SQL.
  • Experience with financial services or payroll data.
  • Expertise in large-scale data processing (billions of records, multi-terabyte data).
  • Knowledge of Delta Lake, Unity Catalog, Databricks Workflows.
  • Experience with CI/CD tools such as Jenkins or Bitbucket Pipelines.
  • Understanding of statistics, data distributions, feature engineering.

Responsibilities

  • Design, develop, and maintain scalable ETL/ELT pipelines for payroll, financial data.
  • Build Databricks-based data platforms using PySpark and Delta Lake.
  • Develop data models and data marts for analytics, reporting, and ML.
  • Optimize data processing across billions of records and multi-terabyte environments.
  • Ensure data quality, lineage, governance, and observability.
  • Translate business requirements into scalable and maintainable technical solutions.
  • Collaborate with data scientists and business stakeholders to deliver data solutions.
  • Support AI-assisted development practices in engineering workflows.

Skills

Databricks
PySpark
Python
SQL

Education

Bachelor's or Master's degree in CS/Data Eng/Info Systems/Statistics/Finance/Economics

Tools

Delta Lake
Unity Catalog
Jenkins
Bitbucket Pipelines
Databricks Asset Bundles

Job description

Position Overview

We are seeking an experienced Lead Data Engineer with strong expertise in Databricks, PySpark, Python, SQL, and modern cloud-based data platforms. The ideal candidate will have a combination of approximately 70% data engineering and 30% analytical responsibilities.

This role involves designing and maintaining scalable data pipelines, optimizing large-scale data processing environments, and supporting financial analytics and data science initiatives.

The successful candidate will bring hands-on experience with financial or payroll data, lakehouse architecture, data governance, and analytical workflows. The individual will collaborate with data scientists, economists, and business stakeholders to deliver reliable data solutions supporting financial research and business intelligence.

Key Responsibilities
Data Engineering & Platform Development
  • Design, develop, and maintain scalable ETL/ELT pipelines for payroll, financial, and macroeconomic datasets.
  • Build and support Databricks-based data platforms using PySpark and Delta Lake.
  • Develop data models and data marts for analytics, reporting, and machine learning use cases.
  • Optimize data processing performance across billions of records and multi-terabyte environments.
  • Ensure data quality, consistency, lineage, governance, and observability.
  • Translate business requirements into scalable and maintainable technical solutions.
Data Analytics & Research Support
  • Perform exploratory data analysis, data profiling, and anomaly detection.
  • Collaborate with data scientists to implement analytical logic using scalable PySpark solutions.
  • Develop validation dashboards and notebooks to verify data quality and pipeline outputs.
  • Support feature engineering, time-series analysis, and complex data aggregations.
  • Utilize Python, pandas, and NumPy for ad hoc analytical tasks.
  • Investigate unusual data patterns and validate analytical results.
Data Architecture & DevOps
  • Work with Databricks, Delta Lake, Unity Catalog, and lakehouse architecture.
  • Participate in architecture discussions involving medallion architecture, catalog design, and data mesh concepts.
  • Implement CI/CD pipelines using Bitbucket Pipelines, Jenkins, and Databricks Asset Bundles.
  • Manage data governance, access controls, and schema design through Unity Catalog.
  • Maintain security, compliance, and data management standards.
AI-Assisted Development
  • Utilize AI-assisted coding tools such as GitHub Copilot, Amazon Q, Kiro, or equivalent.
  • Review AI-generated code for accuracy, performance, scalability, and maintainability.
  • Incorporate AI-assisted development practices into engineering workflows.
Required Qualifications
  • Bachelor's or Master's degree in Computer Science, Data Engineering, Information Systems, Statistics, Finance, Economics, or a related field.
  • 5+ years of experience in Data Engineering or Data Platform Development.
  • Strong hands-on experience with Databricks, PySpark, Python, and SQL.
  • Experience with financial services or payroll data, including compensation, deductions, and pay-period logic.
  • Expertise in large-scale data processing, including billions of records, multi-terabyte datasets, and time-series data.
  • Strong knowledge of Delta Lake, Unity Catalog, Databricks Workflows, and data optimization.
  • Experience with ETL/ELT development, data modeling, and data quality frameworks.
  • Proficiency in pandas and NumPy for exploratory data analysis.
  • Experience with CI/CD tools such as Jenkins, Bitbucket Pipelines, or equivalent.
  • Understanding of statistical concepts, data distributions, correlations, and feature engineering.
  • Experience participating in technical architecture and design decisions.
  • Familiarity with AI-assisted development tools.
Preferred Qualifications
  • Experience with financial markets, macroeconomic, or capital markets datasets.
  • Knowledge of lakehouse architecture, medallion architecture, and data mesh concepts.
  • Experience with Kafka or Spark Structured Streaming.
  • Exposure to Census data, TIGER datasets, and FIPS codes.
  • Infrastructure-as-code experience with Terraform or CDK.
  • Databricks Associate or Professional certification.
  • Experience migrating legacy data platforms to Databricks.
  • Exposure to machine learning and AI-driven analytics.
  • Experience with Power BI, Tableau, or Databricks Dashboards.
  • Scala programming experience is a plus.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Data Engineer
Lead Data Engineer

Compunnel, Inc. • Northern (KY)

On-site
USD 120,000 - 170,000
Sr. Data Engineer
Sr. Data Engineer

Techgene Solutions LLC • Pasadena (CA)

Hybrid
USD 120,000 - 170,000
Senior Data Engineer - Databricks
Senior Data Engineer - Databricks

DATAECONOMY Inc • New Jersey

On-site
USD 130,000 - 180,000
Data Engineer
Data Engineer

Talentify • New York (NY)

On-site
USD 120,000 - 180,000
Databricks Data Engineer
Databricks Data Engineer

Henderson Scott • Irving (TX)

On-site
USD 100,000 - 130,000
Senior Data Engineer - Databricks
Senior Data Engineer - Databricks

DATAECONOMY Inc • Northern (KY)

On-site
USD 140,000 - 180,000
Sr. Data Engineer
Sr. Data Engineer

Techgene Solutions • Pasadena (CA)

Hybrid
USD 120,000 - 180,000
Senior Data Engineer - Databricks
Senior Data Engineer - Databricks

DATAECONOMY Inc • Raleigh (NC)

On-site
USD 120,000 - 190,000
Lead Data Engineer with Databricks
Lead Data Engineer with Databricks

Univedge Consulting LLC • St. Louis (MO)

On-site
USD 120,000 - 180,000
Databricks Engineer
Databricks Engineer

Tredence Inc. • Chicago (IL)

On-site
USD 140,000 - 190,000