Lead Data Engineer (Hands-On)

Brickstech

Cary (NC)

Hybrid

USD 150,000 - 185,000

Full time

20 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Brickstech seeks a true hands-on Lead Data Engineer to own the end-to-end data and AI architecture for an enterprise lakehouse on Azure Databricks. You will guide design, review production code, optimize Spark workloads, and establish engineering guardrails.

The role requires 12–18 years of data engineering leadership, strong Python/Scala skills, and expertise in Delta Lake, Unity Catalog, and modern data QA practices. Hybrid onsite in Cary, NC with collaboration across teams.

Qualifications

  • 12–18 years in data engineering or enterprise data-platform delivery.
  • Proven experience delivering enterprise-scale medallion/lakehouse architectures.
  • Expert-level hands-on experience with Python, Scala, and PySpark.
  • Strong SQL and data modeling skills.
  • Experience with Azure Databricks and Delta Lake.
  • Ability to lead engineers and establish standards.

Responsibilities

  • Own enterprise lakehouse architecture and data contracts.
  • Design scalable, production-grade data pipelines for large workloads.
  • Lead ingestion, streaming and metadata-driven frameworks.
  • Mentor engineers, review code, and enforce testing standards.
  • Drive CI/CD and observability for Databricks and Azure Data Factory.

Skills

Azure Databricks
Python
Scala
PySpark
Data architecture
Production pipelines
Spark tuning
CI/CD pipelines
Leadership
Mentoring

Education

Bachelor's degree in CS or related field

Tools

Azure Data Factory
Databricks Unity Catalog
Delta Lake
Kafka
Event Hubs
Spark Structured Streaming
Terraform
Azure DevOps

Job description

Location: Cary, NC — On-site / HybridEmployment Type: Full-Time
Experience Required: 12–18 years
Eligibility: U.S. Citizens only
About the Engagement
This role supports the development of a centralized, AI-first enterprise Data Hub for a global insurance and financial services organization, built on Azure Databricks.
The platform ingests 150+ inbound data feeds, distributes data to 35+ downstream systems, and follows a modern medallion architecture (Bronze / Silver / Gold).
AI capabilities are embedded throughout the platform from day one, including ingestion, canonical mapping, data quality, reconciliation, anomaly detection, and business-user access.
This is a senior, hands-on technical leadership role. The Lead Data Engineer will own the end-to-end technical design of both the data and AI layers, establish reference implementations for the engineering team, and continue to write and review production-grade Python, Scala, and PySpark code every week.
Candidates who have not written or reviewed production code within the past year will not be considered.
Key Responsibilities
Data Platform Architecture & Engineering

  • Own the enterprise lakehouse architecture, including:
    • Bronze / Silver / Gold data contracts
    • ADLS Gen2 zone design
    • Delta Lake table architecture
    • Partitioning strategies
    • Schema evolution
    • Retention policies
  • Design and implement scalable, production-grade data engineering patterns for enterprise workloads.
Ingestion & Streaming Frameworks
  • Build metadata-driven and parameterized ingestion frameworks supporting:
    • Batch files
    • Database extracts
    • CDC feeds
    • Real-time and streaming workloads
  • Work extensively with:
    • Azure Event Hubs
    • Kafka
    • Spark Structured Streaming
Hands-On Engineering Leadership
  • Develop canonical PySpark and Scala Spark implementations.
  • Establish engineering, coding, and automated testing standards.
  • Conduct code and pull-request reviews.
  • Troubleshoot production incidents and Spark workload failures.
  • Optimize Spark clusters, workloads, and associated cloud costs.
  • Define and enforce engineering guardrails for performance, scalability, and reliability.
CI/CD, Infrastructure & Observability
  • Implement CI/CD pipelines for Databricks and Azure Data Factory using:
    • Azure DevOps
    • Databricks Asset Bundles
    • Terraform
  • Establish observability using:
    • Azure Monitor
    • Log Analytics
  • Drive automated testing and repeatable deployment practices across environments.
AI-Augmented Data Engineering
  • Design AI-assisted ingestion and canonical mapping solutions.
  • Build capabilities for:
    • Automated bridge-document generation
    • DML generation
    • Canonical table-definition generation
    • AI-assisted source-to-canonical mapping
  • Implement appropriate human-review and approval gates for AI-generated mappings and transformations.
AI-Driven Data Quality & Reconciliation
  • Build AI/ML-based solutions for:
    • Data drift detection
    • Schema drift detection
    • Volume anomaly detection
    • Reconciliation failures
    • Automated reconciliation
  • Develop privacy-preserving synthetic test-data approaches.
Semantic Layer & Generative AI
  • Help design the enterprise semantic layer and knowledge graph.
  • Build GPT-powered conversational data-access capabilities, including:
    • Text-to-SQL
    • Semantic-layer retrieval
    • RAG-based access patterns
  • Enforce row-level and column-level security within AI-powered data-access solutions.
Governance & Technical Leadership
  • Implement governance using Databricks Unity Catalog, including:
    • Lineage
    • Access control
    • PII standards
    • Data discovery and governance
  • Participate in and lead:
    • Architecture Review Boards
    • AI governance forums
    • Design reviews
  • Mentor engineers and establish strong technical documentation standards.
  • Present and defend architecture decisions to both technical and non-technical stakeholders.
Must-Have Skills & Experience
Core Data Engineering
  • 12–18 years of total experience in data engineering, data architecture, or enterprise data-platform delivery.
  • Proven experience delivering enterprise-scale medallion / lakehouse architectures.
  • Expert-level hands-on experience with:
    • Python
    • Scala
    • PySpark
  • Strong ability to build modular, well-tested, production-grade data solutions.
  • Deep experience troubleshooting and tuning Spark workloads.
  • Experience optimizing large-scale batch and streaming pipelines using Delta Lake.
SQL & Data Modeling
  • Strong SQL expertise.
  • Strong experience with:
    • Dimensional data modeling
    • Normalized data modeling
    • Schema design
    • Data contracts
Databricks
Strong hands-on experience with:
  • Delta Lake
  • Unity Catalog
  • Databricks Jobs & Workflows
  • Cluster and pool management
  • Performance tuning
  • Databricks Model Serving
Microsoft Azure Data Stack
Strong experience with:
  • ADLS Gen2
    • Zone architecture
    • ACLs
    • Lifecycle management
  • Azure Data Factory
    • Parameterized pipelines
    • Metadata-driven ingestion frameworks
  • Azure Event Hubs
Generative AI / LLM Engineering
  • Minimum 3+ years of experience designing and deploying LLM-based systems in production.
  • Strong hands-on experience with:
    • RAG pipelines
    • Agentic / tool-calling workflows
    • Chunking strategies
    • Embedding strategies
    • Vector retrieval
    • Hybrid retrieval
    • Prompt engineering
  • Experience with at least one of:
    • LangChain
    • LlamaIndex
    • LangGraph
  • Experience with at least one production AI stack:
    • Azure OpenAI
    • OpenAI
    • Databricks Model Serving
LLM Evaluation & Quality
Experience establishing disciplined evaluation processes, including:
  • Golden datasets
  • Regression suites
  • Accuracy measurement
  • Hallucination tracking
  • Human-in-the-loop feedback
Metadata & Governance
Strong experience building or working with metadata-driven frameworks involving:
  • Schema inference
  • Data profiling
  • Lineage
  • Data catalogs
Azure Security
Strong understanding of:
  • Microsoft Entra ID
  • Managed identities
  • RBAC
  • POSIX ACLs
  • Azure Key Vault
  • Private endpoints
  • PII handling and data-security standards
DevOps & Infrastructure as Code
Hands-on experience with:
  • Azure DevOps
  • Terraform
  • Databricks Asset Bundles
  • Automated testing for data pipelines
Communication
  • Strong technical writing skills.
  • Ability to communicate complex architecture decisions clearly.
  • Ability to present and defend technical designs to engineers, leadership, and non-technical stakeholders.
Strongly Preferred
Experience with any of the following is highly desirable:
  • Knowledge graphs and ontologies
    • RDF / SPARQL
    • Neo4j
    • Graph modeling over a lakehouse
  • Enterprise-scale text-to-SQL solutions
  • Semantic-layer-backed natural-language query platforms
  • ML-based anomaly detection for:
    • Time-series data
    • Financial transactions
  • Financial services or insurance domain experience, particularly:
    • Finance close
    • General Ledger
    • Subledger
    • Reconciliation
    • Actuarial data
  • LLMOps / MLOps, including:
    • Model versioning
    • Prompt versioning
    • Cost governance
    • Observability
  • Certifications such as:
    • Databricks Data Engineer Professional
    • Microsoft Azure DP-203
    • Microsoft Fabric DP-700
    • AZ-305
  • Experience with:
    • dbt
    • Great Expectations or similar data-quality platforms
    • Workday
    • Workday Prism
    • Workday Accounting Center
Submission Requirements
Please include the following information:
  • Updated resume
  • Full Name
  • Current Location
  • Contact Number
  • Email Address
  • Work Authorization: U.S. Citizen
  • LinkedIn Profile
  • Availability to Start
  • Interview Availability for the next 3 days
We are looking for a true hands-on Lead Data Engineer who can combine enterprise architecture ownership with day-to-day production engineering. Candidates should be equally comfortable designing the target architecture, reviewing technical decisions, troubleshooting production workloads, and writing production code.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Data Engineer (Hands-On)
Lead Data Engineer (Hands-On)

Cephas Consultancy Services Private Limited • Cary (NC)

Hybrid
USD 150,000 - 210,000
Lead Data Engineer
Lead Data Engineer

Recru • Spring (TX)

On-site
USD 130,000 - 170,000
Lead Data Engineer
Lead Data Engineer

SIDRAM TECHNOLOGIES • Cary (NC)

Hybrid
USD 160,000 - 230,000
Lead Data Engineer – Azure Databricks
Lead Data Engineer – Azure Databricks

Solvimate Pvt. Ltd. • Cary (NC)

Hybrid
USD 140,000 - 145,000
Lead Data Engineer
Lead Data Engineer

K2 Integrity • New York (NY)

On-site
USD 140,000 - 210,000
Data Team Lead
Data Team Lead

Incedo Inc. • San Rafael (CA)

On-site
USD 130,000 - 150,000
Medical insurance
Vision insurance
401(k)
Lead Data Engineer
Lead Data Engineer

Norfolk Southern Corp • Atlanta (GA)

Hybrid
USD 150,000 - 230,000
Sr. Azure Data Engineer
Sr. Azure Data Engineer

Noblesoft Technologies Inc. • Greenfield (IN)

On-site
USD 120,000 - 180,000
Senior Databricks Data Engineer
Senior Databricks Data Engineer

CRED • College Station (TX)

On-site
USD 120,000 - 160,000
Incentive Program
Accrued Time Off
401(k) with match
+4
Senior Data Engineer with Databricks Exp. - 100% Remote
Senior Data Engineer with Databricks Exp. - 100% Remote

SDH Systems • United States

Remote
USD 140,000 - 190,000