Senior Databricks Data Engineer, Spreadsheet & Flat-File Sources

Lumenalta

United States

Remote

USD 120,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Remote work (US)
Career growth

Job summary

Lumenalta seeks a Databricks Data Engineer to design and operate ingestion for Excel and flat-file sources, turning them into governed Delta tables. You will handle schema evolution, data quality, and lineage, joining to BOM and engineering data while documenting every source for client handoff.

The role is fully remote within the United States, requiring US citizenship and collaboration with business owners to pin down definitions and rules.

Qualifications

  • Production experience building ingestion pipelines on Databricks.
  • Hands-on ingestion of Excel and flat-file sources into Spark/Databricks at scale.
  • Experience with schema evolution, data quality, and governed data pipelines.
  • Ability to align ingested data to shared keys for joins with other systems.
  • Strong PySpark and SQL skills; knowledge of Delta Lake and medallion architecture.

Responsibilities

  • Inventory and document Excel workbooks and file exports used by business teams.
  • Build Auto Loader pipelines to ingest workbooks and flat files landed in S3 with explicit schemas.
  • Parse multi-sheet and inconsistently structured files at scale and handle version drift.
  • Configure schema evolution and data quality checks in pipelines.
  • Conform ingested data to keying schemes for joins to BOM and engineering data.
  • Capture lineage, orchestrate with Databricks Workflows, and document sources.

Skills

Databricks
PySpark
SQL
Delta Lake
Auto Loader
Excel ingestion
S3 ingestion
Data quality
Data governance
Documentation

Tools

Databricks Platform

Job description

About the Role

We partner with global enterprises to design and build large-scale data platforms that power the products and operations at the center of their business.

Not all of that data lives in enterprise systems. Manufacturing and supply chain teams maintain parts, supplier, and program data in Excel workbooks and flat-file repositories. These are authoritative sources, but they have no schema, no change log, and no owner-enforced structure. Each workbook is a small, undocumented system of its own.

You will bring those sources onto the platform. You will inventory them, build ingestion that turns inconsistent workbooks and files into governed Delta tables, and make the result join cleanly to parts and revisions from the engineering systems, so the digital thread includes the data that today lives only in spreadsheets.

What You’ll Do
  • Inventory the Excel workbooks and file exports that business and engineering teams rely on, and document their structure, owners, and downstream use.
  • Build Auto Loader pipelines to ingest workbooks and flat files landed in S3, with explicit schemas and typed Bronze and Silver tables.
  • Parse multi-sheet, inconsistently structured workbooks at scale: shifting header rows, merged cells, embedded totals, mixed types in a column, and layouts that differ between versions of the same file.
  • Configure schema evolution in Databricks declarative pipelines, including new column handling, rescued data columns, and controlled failure on renamed columns.
  • Implement data quality expectations that catch the failure modes these sources produce, such as duplicate rows, blank keys, and silently changed formats.
  • Conform ingested data to the platform's part number plus revision keying so it joins to BOM and engineering data from other sources.
  • Ensure lineage is captured for every pipeline, orchestrate with Databricks Workflows, and document each source for handoff.
  • Work with the business owners of each workbook to confirm meaning, business rules, and what "correct" looks like.
  • Translate spreadsheet logic, such as formulas, lookups, and pivot tables, into maintainable SQL and documented business rules.
  • Validate that platform outputs match the numbers users trust today, and explain differences when they do not.
  • Document datasets, metrics, and definitions so reporting stays consistent after handoff.
What We’re Looking For
  • Production experience building ingestion pipelines on Databricks.
  • Hands‑on experience ingesting Excel and flat-file sources into Spark or Databricks at scale and the judgment to know when each is appropriate.
  • A track record of making human‑maintained data reliable: shifting headers, merged cells, mixed types, and layout drift between versions.
  • Experience with Auto Loader for incremental file ingestion from S3.
  • Experience with Databricks declarative pipelines or Delta Live Tables, including schema evolution and data quality expectations.
  • Experience conforming ingested data to shared keys so it joins with data from other systems.
  • Strong PySpark and SQL, and working knowledge of Delta Lake and medallion architecture.
  • Experience working directly with non-technical data owners to pin down definitions and business rules.
  • Ability to complete the client's background check and onboarding and to work on client-furnished equipment.
  • US citizenship.
  • Fluent English, both written and spoken.
Nice to Have
  • Experience extracting from Microsoft Access databases (ODBC, JDBC, or file‑level tooling).
  • Prior work with manufacturing or supply chain data, especially BOMs, parts, revisions, and supplier records.
  • Familiarity with PLM or ERP data, such as Teamcenter or Oracle E‑Business Suite.
  • Prior work in a defense, aerospace, or FedRAMP environment, or with CUI or ITAR‑controlled data.
  • Databricks Data Engineer Professional certification.
Why Lumenalta is an amazing place to work at

At Lumenalta, you can expect that you will:

  • Be 100% dedicated to one project at a time so that you can innovate and grow.
  • Be a part of a team of talented and friendly senior‑level developers.
  • Work on projects that allow you to use leading tech.

For Databricks practitioners specifically, we are a Databricks partner with Champions and MVPs on our team, and that carries three concrete things:

  • Delivery Partner Program qualification. Databricks co‑delivers its Professional Services engagements through a short list of approved partner organizations, and practitioners have to be individually qualified before they can work on them. The qualification is granted once at the individual level and does not expire, so it stays with you. Partner organizations are the route in, so the qualified pool stays small. That scarcity is most of the value, and we sponsor it.
  • Champion path. Databricks Champion is an individual recognition with its own badge, a dedicated Databricks solutions architect as your mentor, access to product teams, confidential roadmap briefings, and an invitation to the internal Tech Summit. Nomination has to come from a partner organization, which puts it out of reach for independents and most contractors. We nominate actively and regularly, and the Champions already on our team mentor candidates through it.
  • Certifications covered. We pay the exam fee for any Databricks certification you want to take.
Our Process

A screening call, a technical interview, and a HackerRank assessment.

The technical interview is a conversation rather than a whiteboard exam. We will ask you to walk through a governance design you have shipped, a federate‑versus‑ingest call you made, what you would change, and what you would expect security reviewers to push hardest on. We will also ask for a writing sample. Relevant project references matter more to us than certifications.

Location

This is a fully remote position.

This position supports a U.S. Government contract that requires all personnel to possess U.S. Citizenship.

This role is remote within the United States only. Candidates must reside in the United States while assigned to our client and be physically located within the United States when performing project work or accessing project systems or data. Availability to work overlapping Pacific, Central, or Eastern U.S. time zones.

Application Deadline

Applications will be accepted until October 4th, 2026. Candidates can expect feedback by October 12th, 2026.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Databricks Data Engineer, PLM & MBSE Ingestion
Senior Databricks Data Engineer, PLM & MBSE Ingestion

Lumenalta • United States

Remote
USD 110,000 - 150,000
Senior Databricks Data Engineer, ERP & Logistics Data
Senior Databricks Data Engineer, ERP & Logistics Data

Lumenalta • United States

Remote
USD 120,000 - 180,000
Senior Databricks Data Engineer, Data Testing & Validation
Senior Databricks Data Engineer, Data Testing & Validation

Lumenalta • United States

Remote
USD 120,000 - 160,000
Senior Databricks Data Engineer, Digital Thread Modeling
Senior Databricks Data Engineer, Digital Thread Modeling

Lumenalta • United States

Remote
USD 120,000 - 180,000
Remote work
Senior Databricks Engineer / Ingestion Lead
Senior Databricks Engineer / Ingestion Lead

Lumenalta • United States

Remote
USD 90,000 - 130,000
Health insurance
401K Contribution
Paid time off
+2
Databricks Data Solution Architect
Databricks Data Solution Architect

Lumenalta • United States

Remote
USD 90,000 - 130,000
Health insurance
401K Contribution
15 days paid time off
+2
Senior Databricks Platform Engineer, Governance & FinOps
Senior Databricks Platform Engineer, Governance & FinOps

Lumenalta • United States

Remote
USD 150,000 - 190,000
Databricks Apps Developer (React)
Databricks Apps Developer (React)

Lumenalta • United States

Remote
USD 120,000 - 160,000
Databricks Data Architect
Databricks Data Architect

Lumenalta • Northern (KY)

Hybrid
USD 90,000 - 130,000
Health insurance
401K contribution
Paid time off
+2
Data Engineer (SAP)
Data Engineer (SAP)

Lumenalta • United States

Remote
USD 120,000 - 180,000