Lead Data Engineer

Compunnel, Inc.

Northern (KY)

Hybrid

USD 120,000 - 170,000

Full time

38 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Compunnel, Inc. is seeking a Lead Data Engineer to drive scalable, cloud-native data platforms for finance analytics.

You will collaborate with data scientists, economists, and business stakeholders to translate complex analytical requirements into production ETL pipelines and robust data models. Responsibilities include building Databricks-based platforms, implementing PySpark/Delta Lake frameworks, ensuring data quality and governance, and shaping lakehouse architecture.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Data Engineering, Economics, Finance or related field.
  • 5+ years of experience in Data Engineering or Data Platform development.
  • Finance or payroll domain experience including payroll data structures and pay-period logic.
  • Experience handling large-scale datasets including billions of records and multi-terabyte environments.
  • Strong proficiency with Databricks including Unity Catalog, Delta Lake, and Workflows.
  • Experience with Python, SQL, PySpark, data modeling, and ETL/ELT development.
  • Analytical fluency with EDA, distributions, time-series patterns, and feature engineering.
  • Experience implementing CI/CD using Bitbucket Pipelines, Jenkins, and automated deployment frameworks.
  • Experience with AI-assisted development tools such as GitHub Copilot, Amazon Q, Kiro, or equivalents.
  • Experience with data governance, data quality, and metadata management.

Responsibilities

  • Design, develop, and maintain scalable data pipelines for ingestion, transformation, and distribution of payroll, macroeconomic, and financial datasets.
  • Build and support Databricks-based platforms enabling financial research and analytical workloads.
  • Implement ETL/ELT using PySpark and Delta Lake for structured and unstructured data.
  • Develop data models and data marts for analytical, reporting, and ML use cases.
  • Ensure data quality, lineage, governance, and observability across data assets.
  • Optimize performance for large-scale datasets
  • Translate business requirements into scalable data solutions.
  • Perform exploratory data analysis and dataset profiling.
  • Translate data scientist logic into efficient PySpark implementations.
  • Build validation dashboards and notebooks to verify outputs and quality.
  • Support feature engineering by implementing complex aggregation and transformation logic at scale.
  • Independently validate analytical outputs and conduct sanity checks.
  • Conduct ad hoc analytics using pandas and NumPy alongside PySpark.
  • Contribute to lakehouse architecture discussions and data mesh concepts.
  • Implement CI/CD pipelines using Databricks Asset Bundles, Bitbucket Pipelines, Jenkins.

Skills

PySpark
Python
SQL
Databricks
Delta Lake
Data modeling
large-scale data processing
finance/payroll domain
CI/CD
GitHub Copilot / AI-assisted coding

Education

Bachelor's or Master's in CS/Data Eng/related

Tools

Unity Catalog
Databricks Asset Bundles
Bitbucket Pipelines
Jenkins
Pandas/NumPy
Tableau/Power BI

Job description

The Lead Data Engineer will work across data engineering and analytical responsibilities, with a primary focus on building scalable cloud-native data platforms and supporting financial analytics and research. The role requires strong expertise in Databricks, PySpark, Python, SQL, Delta Lake, data modeling, and large-scale data processing, along with finance or payroll domain experience. The engineer will collaborate closely with data scientists, economists, and business stakeholders to develop production ETL pipelines, explore and profile datasets, optimize data platforms, and translate complex analytical requirements into scalable solutions.

Key Responsibilities
  • Design, develop, and maintain scalable data pipelines for the ingestion, transformation, and distribution of payroll, macroeconomic, and financial datasets.
  • Build and support Databricks-based platforms that enable financial research and analytical workloads.
  • Implement ETL/ELT frameworks using PySpark and Delta Lake for structured and unstructured data from internal and external sources.
  • Develop data models and data marts optimized for analytical, reporting, and machine learning use cases.
  • Ensure data quality, consistency, lineage, governance, and observability across data assets.
  • Optimize performance for large-scale datasets, including billions of records, multi-terabyte environments, and time-series data.
  • Translate business requirements into scalable and maintainable data solutions.
  • Perform exploratory data analysis, including dataset profiling and identification of distributions, outliers, missing patterns, and data drift.
  • Translate data scientist logic into efficient and scalable PySpark implementations, including cross-sectional metrics and time-windowed aggregations.
  • Build validation dashboards and exploratory notebooks to verify pipeline outputs and data quality.
  • Support feature engineering by implementing complex aggregation and transformation logic at scale.
  • Independently validate analytical outputs, perform sanity checks, and identify results that warrant further investigation.
  • Conduct ad hoc analytical work using pandas and NumPy alongside PySpark to support research initiatives.
  • Contribute to lakehouse architecture design discussions and evaluate tradeoffs related to catalog design, medallion architecture, and data mesh concepts.
  • Implement CI/CD pipelines using Databricks Asset Bundles, Bitbucket Pipelines, Jenkins, and automated deployment frameworks.
  • Manage Unity Catalog governance, access patterns, and schema design.
  • Ensure security, compliance, and data governance standards are maintained.
  • Leverage AI-assisted coding tools such as GitHub Copilot, Amazon Q, Kiro, or equivalent platforms to accelerate development.
  • Review AI-generated code for correctness, performance, scalability, and maintainability.
  • Integrate AI-assisted development workflows into engineering and analytical activities.
Required Qualifications
  • Bachelor's or Master's degree in Computer Science, Data Engineering, Information Systems, Statistics, Economics, Finance, or a related field.
  • 5+ years of experience in Data Engineering or Data Platform development.
  • Finance or payroll domain experience, including familiarity with payroll data structures, pay-period logic, compensation and deduction relationships, or financial-services data environments.
  • Experience handling large-scale datasets, including billions of records, multi-terabyte environments, and time-series data.
  • Strong proficiency with Databricks, including Unity Catalog, Delta Lake, Databricks Workflows, and Databricks Asset Bundles or equivalent deployment frameworks.
  • Strong understanding of Unity Catalog governance, access patterns, and catalog/schema design.
  • Strong understanding of Delta Lake internals, including optimization, clustering, change data feed, and versioning.
  • Experience with Databricks Workflows, including orchestration, dependencies, and monitoring.
  • Experience participating in architecture-level design decisions and evaluating technical tradeoffs.
  • Strong proficiency in Python, SQL, PySpark, data modeling, and ETL/ELT development.
  • Analytical fluency with exploratory data analysis, basic statistical concepts, distributions, correlations, time-series patterns, and feature engineering.
  • Proficiency with pandas and NumPy for ad hoc analytical work alongside production PySpark.
  • Experience implementing CI/CD using tools such as Bitbucket Pipelines, Jenkins, and automated deployment frameworks.
  • Experience with AI-assisted development tools such as GitHub Copilot, Amazon Q, Kiro, or equivalent.
  • Experience implementing data quality and validation frameworks.
  • Strong understanding of data governance, data quality, and metadata management.
  • Strong analytical and problem-solving skills with statistical literacy.
  • Excellent communication and documentation skills.
  • Ability to work effectively in a fast-paced, data-driven environment and collaborate with technical and business stakeholders.
Preferred Qualifications
  • Experience with macroeconomic, capital markets, or financial-services data.
  • Exposure to lakehouse patterns, data mesh concepts, and medallion architecture.
  • Experience with streaming and event-driven pipelines using Kafka or Structured Streaming.
  • Census or geographic data processing experience, including TIGER and FIPS codes.
  • Infrastructure-as-code experience with Terraform, CDK, or similar technologies.
  • Experience migrating legacy data platforms to modern technology stacks, including Glue, EMR, or HDInsight to Databricks.
  • Experience supporting machine learning and AI-driven analytics solutions.
  • Data visualization experience using Power BI, Tableau, Databricks Dashboards, or similar platforms.
  • Experience with Python visualization libraries such as Matplotlib and Plotly.
  • Experience with SQL Server, PostgreSQL, Delta Tables, or NoSQL databases.
  • Experience with Scala.
  • Experience with large-scale time-series data engineering.
  • Experience with payroll data structures and processing.
  • Experience with macroeconomic analysis and forecasting.
  • Experience with financial markets and alternative data.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Databricks Engineer
Databricks Engineer

iLink Digital • Milpitas (CA), Northern (KY)

On-site
USD 120,000 - 160,000
Data Engineering Architect, Senior
Data Engineering Architect, Senior

Bloomberg • Virginia (MN)

On-site
USD 120,000 - 160,000
Data Engineer
Data Engineer

Ranger Technical Resources • Town of Florida (NY)

On-site
USD 130,000 - 185,000
Data Engineer
Data Engineer

Recru, LLC. • Spring (TX)

On-site
USD 180,000 - 230,000
Data Engineer
Data Engineer

Scorpion Therapeutics • Indianapolis (IN)

On-site
USD 120,000 - 180,000
Lead Data Engineer
Lead Data Engineer

MetLife • Bridgewater (MA)

On-site
USD 140,000 - 185,000
Databricks Technical Lead
Databricks Technical Lead

Anblicks • Dallas (TX)

On-site
USD 120,000 - 160,000
Lead Software Engineer - Databricks
Lead Software Engineer - Databricks

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 150,000 - 190,000
BI Data Engineer II
BI Data Engineer II

Jobtailor • Boston (MA)

On-site
USD 120,000 - 180,000
Lead Data Engineer
Lead Data Engineer

K2 Intelligence, LLC • Northern (KY)

Hybrid
USD 120,000 - 180,000