Senior Databricks Data Engineer, Data Testing & Validation

Lumenalta

United States

Remote

USD 120,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Lumenalta is seeking a data engineer to design and implement automated tests for its large-scale data platforms. You will write reconciliation tests between Databricks tables and legacy reports, validate incremental loads, and ensure idempotent re-runs while preventing duplicates across sources.

Responsibilities include building PySpark/SQL test frameworks, integrating tests into Databricks Workflows and CI/CD, and defining data quality gates.

Qualifications

  • 5+ years in data engineering, including building test/validation frameworks for production pipelines.
  • Strong PySpark and SQL, used to write reconciliation and assertion logic at scale.
  • Hands-on experience reconciling a new platform against a legacy system or its reports.
  • Deep understanding of incremental load correctness: CDC, idempotency, MERGE, late-arriving data.
  • Experience with Python test frameworks like pytest and data testing tools (chispa, Great Expectations).
  • Experience with Databricks Workflows and CI/CD across environments.
  • Clear written communication; plan and report test evidence and variances.

Responsibilities

  • Build the automated test framework for pipelines in PySpark and SQL and run it across environments.
  • Write reconciliation tests comparing Databricks tables to legacy systems and reports with clear variance output.
  • Test incremental loading, ensure idempotency, correct merges/upserts, and no double-loads.
  • Check duplicates on composite keys and cross-source relationships.
  • Write unit tests for transformation logic to validate keys, joins, and business rules.
  • Define data quality checks within pipelines to quarantine bad data.
  • Test behavior when source files change shape (new/renamed/missing columns).
  • Integrate tests with Databricks Workflows and the CI/CD pipeline for gatekeeping deployments.
  • Verify Unity Catalog access controls, row filters, and data masking for restricted data.
  • Document test results and defects, and present weekly quality status.

Skills

Data engineering
PySpark
SQL
Data validation
Unit testing
Delta Lake
CI/CD
Databricks Workflows
US citizenship
Clear written communication

Tools

Databricks
Great Expectations
chispa
Delta Live Tables
Lakewood?

Job description

About the Role

We partner with global enterprises to design and build large-scale data platforms that power the products and operations at the center of their business.

The platform replaces an existing integration system that today connects the client's engineering tools and produces the reports their teams rely on. The rule is simple: the data in Databricks has to match what that system holds and what its reports say. It has to match on day one, and it has to keep matching every time new data arrives.

You will write the tests that prove it. You are a data engineer, not a manual tester. Your code checks that every table reconciles to the legacy system or its reports, that incremental loads add exactly what changed and nothing twice, and that duplicates never make it through. When a pipeline breaks one of those rules, your tests catch it before the client does.

What You'll Do
  • Build the automated test framework for the platform's pipelines, in PySpark and SQL, and run it on every deployment throughout all environments.
  • Write reconciliation tests that compare Databricks tables to the legacy system and its reports, with clear output on where they differ and why.
  • Test incremental loading: change data capture, appends, and batch refreshes. Prove that reruns are idempotent, that merges and upserts update the right rows, and that late or re-delivered files do not double-load.
  • Test for duplicates on composite keys and on the relationships that link records across source systems.
  • Write unit tests for transformation logic so that keying, joins, and business rules are verified in isolation before they run against real data.
  • Define data quality expectations in the pipelines themselves, so bad data is quarantined or fails the run instead of landing silently.
  • Test how pipelines behave when source files change shape: new columns, renamed columns, missing columns, and mixed types.
  • Wire the tests into Databricks Workflows and the CI/CD pipeline so they gate promotion between environments.
  • Test Unity Catalog access controls, row filters, column masks, and classification tags to confirm that export-controlled data is only visible to authorized users.
  • Plan and run user acceptance testing sessions with client stakeholders, capture feedback, and track defects to resolution.
  • Document test results and acceptance evidence for each deliverable, and report quality status in weekly reviews.
What We're Looking For
  • 5+ years in data engineering, including building test and validation frameworks for production pipelines, not only running tests written by others.
  • Strong PySpark and SQL, used to write reconciliation and assertion logic at scale.
  • Hands‑on experience reconciling a new platform against a legacy system or its reports, and explaining the variances.
  • Deep understanding of incremental load correctness: CDC and append semantics, idempotency, MERGE behavior, watermarks, and late-arriving data.
  • Experience finding and resolving duplicates on composite keys across multiple sources.
  • Experience with a Python test framework such as pytest, and with data testing tools such as chispa, Great Expectations, or pipeline expectations in Delta Live Tables or Lakeflow.
  • Experience running tests from Databricks Workflows and CI/CD across multiple environments.
  • Organized, detail-oriented, and comfortable holding a delivery team and a client to a quality bar on a tight timeline.
  • Clear written communication - test evidence and variance reports are the core deliverables.
  • Ability to complete the client's background check and onboarding and to work on client-furnished equipment.
  • US citizenship.
  • Fluent English, both written and spoken.
Nice to Have
  • Hands‑on experience with Databricks, Unity Catalog, or Delta Lake.
  • Python or PySpark for automating data validation.
  • Testing Unity Catalog access controls, row filters, and column masks.
  • Prior work in a defense, aerospace, or FedRAMP environment, or with CUI or ITAR-controlled data.
  • Familiarity with PLM or MBSE tools such as Teamcenter or Cameo, or with Oracle E-Business Suite data.
  • Experience with test management and defect tracking tools such as Jira or Xray.
Why Lumenalta is an amazing place to work at

At Lumenalta, you can expect that you will:

  • Be 100% dedicated to one project at a time so that you can innovate and grow.
  • Be a part of a team of talented and friendly senior-level developers.
  • Work on projects that allow you to use leading tech.

For Databricks practitioners specifically, we are a Databricks partner with Champions and MVPs on our team, and that carries three concrete things:

  • Delivery Partner Program qualification. Databricks co-delivers its Professional Services engagements through a short list of approved partner organizations, and practitioners have to be individually qualified before they can work on them. The qualification is granted once at the individual level and does not expire, so it stays with you. Partner organizations are the route in, so the qualified pool stays small. That scarcity is most of the value, and we sponsor it.
  • Champion path. Databricks Champion is an individual recognition with its own badge, a dedicated Databricks solutions architect as your mentor, access to product teams, confidential roadmap briefings, and an invitation to the internal Tech Summit. Nomination has to come from a partner organization, which puts it out of reach for independents and most contractors. We nominate actively and regularly, and the Champions already on our team mentor candidates through it.
  • Certifications covered. We pay the exam fee for any Databricks certification you want to take.
Our Process

A screening call, a technical interview, and a HackerRank assessment.

The technical interview is a conversation rather than a whiteboard exam. We will ask you to walk through an Oracle integration you built, how you handled security and access constraints, how you modeled logistics or ERP data alongside engineering data, and what you would change. Relevant project experience matters more to us than certifications.

Location

This is a fully remote position.

This position supports a U.S. Government contract that requires all personnel to possess U.S. Citizenship.

This role is remote within the United States only. Candidates must reside in the United States while assigned to our client and be physically located within the United States when performing project work or accessing project systems or data. Availability to work overlapping Pacific, Central, or Eastern U.S. time zones.

Application Deadline

Applications will be accepted until October 4th, 2026. Candidates can expect feedback by October 12th, 2026.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Databricks Data Engineer, Spreadsheet & Flat-File Sources
Senior Databricks Data Engineer, Spreadsheet & Flat-File Sources

Lumenalta • United States

Remote
USD 120,000 - 160,000
Remote work (US)
Career growth
Senior Databricks Data Engineer, PLM & MBSE Ingestion
Senior Databricks Data Engineer, PLM & MBSE Ingestion

Lumenalta • United States

Remote
USD 110,000 - 150,000
Senior Databricks Data Engineer, ERP & Logistics Data
Senior Databricks Data Engineer, ERP & Logistics Data

Lumenalta • United States

Remote
USD 120,000 - 180,000
Senior Databricks Platform Engineer, Governance & FinOps
Senior Databricks Platform Engineer, Governance & FinOps

Lumenalta • United States

Remote
USD 150,000 - 190,000
Senior Databricks Data Engineer, Digital Thread Modeling
Senior Databricks Data Engineer, Digital Thread Modeling

Lumenalta • United States

Remote
USD 120,000 - 180,000
Remote work
Senior Databricks Engineer / Ingestion Lead
Senior Databricks Engineer / Ingestion Lead

Lumenalta • United States

Remote
USD 90,000 - 130,000
Health insurance
401K Contribution
Paid time off
+2
Databricks Data Solution Architect
Databricks Data Solution Architect

Lumenalta • United States

Remote
USD 90,000 - 130,000
Health insurance
401K Contribution
15 days paid time off
+2
Databricks Apps Developer (React)
Databricks Apps Developer (React)

Lumenalta • United States

Remote
USD 120,000 - 160,000
Databricks Data Architect
Databricks Data Architect

Lumenalta • Northern (KY)

Hybrid
USD 90,000 - 130,000
Health insurance
401K contribution
Paid time off
+2
Business Architect, Databricks (Healthcare/Clinical Operations Business Lead)
Business Architect, Databricks (Healthcare/Clinical Operations Business Lead)

Lumenalta • United States

Remote
USD 130,000 - 170,000