Senior Databricks Engineer / Ingestion Lead

Lumenalta

United States

Remote

USD 90,000 - 130,000

Full time

11 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
401K Contribution
Paid time off
Professional development fund
Remote work

Job summary

Lumenalta is seeking a Senior Databricks Engineer to own the ingestion layer for a US defense client. You will run source discovery across PLM, ERP, and MBSE sources, design ingestion patterns (Lakeflow Connect, JDBC, Auto Loader) and build end-to-end pipelines with careful handling of schema evolution.

This fully remote role requires US citizenship, offers a salary range of $90,000-$130,000, strong benefits, and a path to Databricks certifications plus working with leading teams on complex data

Qualifications

  • 7+ years of production data engineering on Databricks.
  • Strong Python, PySpark, and SQL.
  • Fluent in Delta Lake, Auto Loader, and declarative pipeline development.
  • Hands-on experience with Lakeflow Connect.
  • Direct experience building JDBC-based ingestion against Oracle E-Business Suite, and familiarity with its security model, session context, and multi-org data structures.
  • Experience with SQL Server as an ingestion source.
  • Working fluency in manufacturing PLM and MBSE systems: Teamcenter, Cameo, and Siemens Capital.
  • Databricks Data Engineer Professional certification.
  • US citizenship and ability to pass a background check.
  • Fluent English, both written and spoken, with strong collaboration skills in distributed teams.

Responsibilities

  • Run source discovery across each system, documenting data structures, change patterns, extraction constraints, data quality, and current-state integration topology.
  • Assess extraction options per source, including change data capture feasibility from PLM, JDBC access to Oracle E-Business Suite under its row-level security model, SQL Server change patterns, and file-based ingestion from S3.
  • Design the ingestion pattern for each source (Lakeflow Connect, custom JDBC, Auto Loader, or federation) with the tradeoffs written down and defensible.
  • Model bill-of-materials data across PLM, ERP, and MBSE sources, keyed on part number plus revision, including reconciliation into bronze and silver layers.
  • Build the ingestion layer: Lakeflow Connect pipelines including CDC from Teamcenter PLM, JDBC connectors against Oracle E-Business Suite with session parameter handling, Auto Loader for incremental S3 ingestion, and SQL Server ingestion patterns.
  • Handle schema evolution deliberately, with automatic handling of new columns, rescued data columns for unexpected shapes, and controlled failure on renamed columns.
  • Operate nightly batch patterns with idempotent reruns and backfill support, quarantining bad source data rather than dropping it.
  • Apply strong engineering practices including testing, CI/CD, and version control, and write expectations and quality checks into pipelines.
  • Document for handoff, producing each source's ingestion contract and runbook so the client team can own it.

Skills

Databricks
Python
PySpark
SQL
Delta Lake
Auto Loader
CI/CD
Git workflows

Tools

Lakeflow Connect
Oracle E-Business Suite
SQL Server
Terraform
Unity Catalog

Job description

We partner with global enterprises to design and build large-scale data platforms that power the products and operations at the center of their business. This engagement is a Databricks lakehouse foundation for a US defense manufacturer, and we are looking for a Senior Databricks Engineer to own the ingestion layer.

The sources are the job. Change data capture out of a PLM system. JDBC ingestion against an ERP where session-level initialization parameters decide whether row-level security hands you the right rows or quietly the wrong ones. Bill-of-materials reconciliation across three systems of record that each describe the same part differently, where someone has to decide which is right.

This is a hands-on senior role with real ownership. Two people on the engagement means ingestion is yours end-to-end: you make the pattern calls per source and sit in the client discovery sessions rather than receiving tickets from them, working alongside the engagement architect who owns platform direction and governance.

The engagement opens with a discovery and design phase, where you go deep on each source system and design the right ingestion pattern for it. It then moves into the build, executing what you designed.

Actively hiring

We are hiring for a current opening on an active client project. This is a specific, presently open role. We review applications on a rolling basis and aim to move qualified candidates through our process promptly.

What You'll Do
  • Run source discovery across each system in scope, documenting data structures, change patterns, extraction constraints, data quality, and current-state integration topology.
  • Assess extraction options per source, including change data capture feasibility from PLM, JDBC access to Oracle E-Business Suite under its row-level security model, SQL Server change patterns, and file-based ingestion from S3.
  • Design the ingestion pattern for each source (Lakeflow Connect, custom JDBC, Auto Loader, or federation) with the tradeoffs written down and defensible.
  • Model bill-of-materials data across PLM, ERP, and MBSE sources, keyed on part number plus revision, including how conflicting representations of the same part are reconciled into the bronze and silver layers.
  • Build the ingestion layer: Lakeflow Connect pipelines including change data capture from Teamcenter PLM, custom JDBC connectors against Oracle E-Business Suite with session-level initialization parameter handling, Auto Loader for incremental S3 ingestion, and SQL Server ingestion patterns.
  • Handle schema evolution deliberately, with automatic handling of new columns, rescued data columns for unexpected shapes, and controlled failure on renamed columns rather than silent data loss.
  • Operate nightly batch patterns with idempotent reruns and backfill support, quarantining bad source data rather than dropping it.
  • Apply strong engineering practices including testing, CI/CD, and version control, and write expectations and quality checks into pipelines so failures are attributable to a source and a rule.
  • Document for handoff, producing each source's ingestion contract and runbook so the client team can own it.
What We're Looking For
  • 7+ years of production data engineering on Databricks, with strong Python, PySpark, and SQL. You are fluent in Delta Lake, Auto Loader, and declarative pipeline development.
  • Hands-on experience with Lakeflow Connect or equivalent managed change data capture ingestion.
  • Direct experience building JDBC-based ingestion against Oracle, ideally Oracle E-Business Suite, and familiarity with its security model, session context, and multi-org data structures.
  • Experience with SQL Server as an ingestion source.
  • Working fluency in manufacturing PLM and MBSE systems: Teamcenter, Cameo, and Siemens Capital. You do not need to be an administrator, but you need to understand their data models well enough to design extraction correctly.
  • Experience modeling bill-of-materials data, including revision handling, effectivity, and multi-source reconciliation.
  • Solid experience with Git-based workflows, CI/CD pipelines, and testing frameworks such as PyTest.
  • Strong AWS experience, particularly S3 and IAM roles.
  • Databricks Data Engineer Professional certification.
  • Experience in a defense, aerospace, or otherwise regulated data environment.
  • Ability to do discovery work: sit with a source system owner, ask the right questions, and document accurately what you find.
  • A track record of building pipelines that survive contact with messy real-world sources and that other engineers can maintain.
  • US citizenship and ability to pass a background check.
  • Fluent English, both written and spoken, with strong collaboration skills in distributed teams.
Nice to Have
  • Prior Teamcenter or other PLM database extraction experience, including direct schema knowledge.
  • Familiarity with Unity Catalog governance from the producer side: tagging, documentation, and lineage-friendly pipeline design.
  • Experience with Infrastructure as Code using Terraform, and with Databricks Asset Bundles.
  • Experience integrating AI-native tooling into engineering workflows.
Why Lumenalta is an amazing place to work

At Lumenalta, you can expect that you will:

  • Be 100% dedicated to one project at a time so that you can innovate and grow.
  • Be a part of a team of talented and friendly senior-level developers.
  • Work on projects that allow you to use leading tech. We have built for Bloomberg, Target, and Two Sigma, and data platform work for Redwood Logistics and Arbol.

For Databricks practitioners specifically, we are a Databricks partner with Champions and MVPs on our team, and that carries three concrete things:

  • Delivery Partner Program qualification. Databricks co-delivers its Professional Services engagements through a short list of approved partner organizations, and practitioners have to be individually qualified before they can work on them. The qualification is granted once at the individual level and does not expire, so it stays with you. Partner organizations are the route in, so the qualified pool stays small. That scarcity is most of the value, and we sponsor it.
  • Champion path. Databricks Champion is an individual recognition with its own badge, a dedicated Databricks solutions architect as your mentor, access to product teams, confidential roadmap briefings, and an invitation to the internal Tech Summit. Nomination has to come from a partner organization, which puts it out of reach for independents and most contractors. We nominate actively and regularly, and engineers are nominated here as readily as architects: someone with hard-won extraction knowledge in systems most people never touch has more to contribute to that community than most.
  • Certifications covered. We offer the exam fee for any Databricks certification, as we are official partners.
Our Process

A screening call, a technical interview, and a HackerRank assessment.

The technical interview is a conversation rather than a whiteboard exam. We will ask you to walk through an ingestion pipeline you built against a difficult source, how you handled schema drift and bad data, and how you would approach discovery and change data capture design for a source system nobody has documented. Relevant project references and specific source system experience carry more weight with us than certifications.

Salary range: $90,000 - $130,000 annually, with final compensation determined by your qualifications, expertise, experience, and the role's scope.

Location:

This is a fully remote position. This position supports a U.S. Government contract that requires all personnel to possess U.S. Citizenship. This role is remote within the United States only. Candidates must reside in the United States while assigned to our client and be physically located within the United States when performing project work or accessing project systems or data. Availability to work overlapping Pacific, Central, or Eastern U.S. time zones.

In addition to competitive pay, we offer a variety of benefits to support your professional and personal growth, including:

  • Flexible working hours in a remote environment.
  • Health insurance (medical and dental) for W2 Employees.
  • 401K Contribution.
  • A professional development fund to enhance your skills and knowledge.
  • 15 days of paid time off annually.
  • Access to soft-skill development courses to further your career.

This is a full-time position requiring a minimum of 40 hours per week, Monday through Friday.

At Lumenalta, we are committed to creating an environment that prioritizes growth, work-life balance, and the diverse needs of our team members.

Application Deadline

Applications will be accepted until October 4th, 2026. Candidates can expect feedback by October 12th, 2026.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Databricks Data Solution Architect
Databricks Data Solution Architect

Lumenalta • United States

Remote
USD 90,000 - 130,000
Health insurance
401K Contribution
15 days paid time off
+2
Senior Databricks Engineer / Ingestion Lead
Senior Databricks Engineer / Ingestion Lead

Lumenalta (formerly Clevertech) • Northern (KY)

On-site
USD 90,000 - 130,000
Health insurance
Remote work environment
Professional development fund
+1
Databricks Data Architect
Databricks Data Architect

Lumenalta • Northern (KY)

Hybrid
USD 90,000 - 130,000
Health insurance
401K contribution
Paid time off
+2
AWS DevOps Engineer - Tech Lead
AWS DevOps Engineer - Tech Lead

Lumenalta • United States

Remote
USD 65,000 - 130,000
Health insurance
401K Contribution
Professional development fund
+2
Solutions Architect - Databricks (Remote)
Solutions Architect - Databricks (Remote)

Lumenalta • Dallas (TX)

On-site
USD 75,000 - 140,000
Flexible working hours
Health insurance
401K Contribution
+3
Data Solutions Architect / Technical Lead
Data Solutions Architect / Technical Lead

Lumenalta • United States

Remote
CAD 104,000 - 222,000
Fully remote position
Data Engineer - Databricks
Data Engineer - Databricks

Lumenalta • Salt Lake City (UT)

On-site
USD 60,000 - 110,000
Flexible working hours
Health insurance
401(k) contribution
+3
Lead Data Engineer - Databricks (Remote)
Lead Data Engineer - Databricks (Remote)

Lumenalta • Raleigh (NC)

On-site
USD 65,000 - 130,000
Flexible working hours
Health insurance (medical and dental)
401K Contribution
+3
Data Engineer - Databricks (Remote)
Data Engineer - Databricks (Remote)

Lumenalta • Raleigh (NC)

On-site
USD 60,000 - 110,000
Flexible working hours
Health insurance (medical and dental)
401K Contribution
+3
Senior DataBricks Developer (Healthcare)
Senior DataBricks Developer (Healthcare)

Lumenalta • Chicago (IL)

On-site
USD 60,000 - 120,000
Flexible working hours
Health insurance (medical and dental)
401K Contribution
+3