Lead Data Engineer (Hands-On)

Cephas Consultancy Services Private Limited

Cary (NC)

Hybrid

USD 150,000 - 210,000

Full time

18 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Cephas Consultancy Services Private Limited in Cary, NC seeks a senior Lead Data Engineer to own end-to-end data and AI layer design on Azure Databricks. You will ship production-grade Python, Scala, and PySpark code weekly and lead data platform initiatives across ingestion, governance, and semantic layers.

The role requires extensive hands-on experience with Delta Lake, Unity Catalog, and Azure data tools, plus leadership of engineering teams in a fast-paced environment.

Qualifications

  • Extensive production-grade experience with Python, Scala and PySpark.
  • Strong SQL/data modeling skills including dimensional modelling.
  • Hands-on Databricks with Delta Lake and Unity Catalog; CI/CD for data pipelines.
  • Azure data stack expertise and experience with ADF, ADLS Gen2, Event Hubs.
  • 3+ years building LLM-based systems in production; ML techniques for data apps.

Responsibilities

  • Own end-to-end technical design of data and AI layers in a medallion/lakehouse setup.
  • Deliver production-grade PySpark/Scala code weekly and review others' code.
  • Architect data platform: lakehouse layout, partitioning, schema evolution, retention.
  • Lead parameterized ingestion for batch, CDC, and streaming sources; ensure quality.
  • Implement CI/CD for Databricks and Azure DevOps; monitor observability.
  • Governance: Unity Catalog, RBAC, and security practices; mentor engineers.

Skills

Python
Scala
PySpark
SQL
Data Modeling
Databricks
Azure Data Stack
LLMs
LangChain
CI/CD
IaC
Technical Writing

Tools

Delta Lake
Unity Catalog
Azure DevOps
Terraform
Databricks Asset Bundles
Azure Data Factory

Job description

Lead Data Engineer (Hands-On)

Cephas Consultancy Services Private Limited Cary, North Carolina, United States

About this position

Client:

Employment: Full-Time

(On-site / Hybrid)

Salary: $ - $ per annum plus benefits
Eligibility: US Citizens only

ABOUT THE ENGAGEMENT

A centralized, AI-first enterprise Data Hub for a global insurance and financial services client on Azure Databricks. The platform ingests 150+ inbound data feeds, distributes to 35+ downstream systems, and is organized as a medallion architecture (Bronze / Silver / Gold). AI is embedded in ingestion, canonical mapping, data quality, reconciliation, and business user access from day one.

This is a senior hands-on leadership role. The candidate will own the end-to-end technical design of the data and AI layers, build reference implementations for the engineering team, and ship production-grade Python, Scala, and PySpark code every week. Candidates who have not written or reviewed production code in the past year are not a fit.

WHAT THE ROLE OWNS
  • Data platform architecture and engineering: lakehouse architecture (Bronze / Silver / Gold contracts, ADLS Gen2 zone layout, Delta Lake table design, partitioning, schema evolution, retention).
  • Metadata-driven, parameterized ingestion frameworks for batch files, database extracts, CDC feeds and streaming (Azure Event Hubs / Kafka,
    Spark Structured Streaming).
  • Canonical PySpark and Scala Spark jobs, coding and testing standards, PR reviews, production incident debugging, Spark cluster tuning and cost guardrails.
  • CI/CD for Databricks and ADF in Azure DevOps using Databricks Asset
    Bundles and Terraform; observability with Azure Monitor and Log
    Analytics.
  • AI-augmented ingestion and canonical mapping: auto-generated bridge documents, DML, canonical table definitions; AI-assisted source-to-canonical mapping with human review gate.
  • AI-driven data quality, anomaly detection (data drift, schema drift,
    volume shifts, reconciliation breaks), automated reconciliation, and
    synthetic privacy-preserving test data.
  • Semantic layer and knowledge graph, plus a GPT-powered
    conversational interface (text-to-SQL / semantic-layer retrieval) with
    row- and column-level security.
  • Governance and leadership: Unity Catalog (lineage, access control,
    PII standards), Architecture Review Boards and AI governance forums,
    mentoring engineers, documentation.
MUST-HAVE SKILLS & EXPERIENCE
  • Expert-level Python, Scala and PySpark: production-ready, modular,
    well-tested solutions; Spark workload troubleshooting; optimizing
    large-scale batch and streaming pipelines using Delta Lake.
  • Strong SQL and data modelling (dimensional and normalised), schema
    design, data contracts.
  • Databricks expertise: Delta Lake, Unity Catalog, Jobs & Workflows,
    cluster and pool management, performance tuning, Model Serving.
  • Azure data stack: ADLS Gen2 (zone design, ACLs, lifecycle), Azure
    Data Factory (parameterized / metadata-driven frameworks), Azure Event
    Hubs.
  • 3+ years designing and shipping LLM-based systems in production: RAG
    pipelines, agentic / tool-calling workflows, chunking and embedding
    strategy, vector and hybrid retrieval, prompt engineering.
  • Evaluation discipline: golden datasets, regression suites, accuracy
    and hallucination tracking, human-in-the-loop feedback.
  • Hands-on with LangChain, LlamaIndex or LangGraph, plus at least one
    provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
  • Metadata-driven frameworks: schema inference, data profiling,
    lineage, catalogs.
  • 12-18 years of total experience in data engineering / data platform delivery.
  • Proven enterprise-scale delivery of a medallion / lakehouse architecture.
  • Azure security and governance: Entra ID, managed identities, RBAC,
    POSIX ACLs, Key Vault, private endpoints, PII handling.
  • CI/CD and IaC: Azure DevOps, Terraform, Databricks Asset Bundles,
    automated testing of data pipelines.
  • Clear technical writing and ability to present and defend designs to
    engineers and non-technical stakeholders.
STRONGLY PREFERRED
  • Knowledge graphs and ontologies (RDF/SPARQL, Neo4j, graph modelling
    over a lakehouse).
  • Text-to-SQL or semantic-layer-backed natural-language query systems
    at enterprise scale.
  • ML-based anomaly detection on time-series or transactional financial data.
  • Financial services or insurance domain (finance close, GL,
    subledger, reconciliation, actuarial data).
  • LLMOps / MLOps: model and prompt versioning, cost governance, observability.
  • Databricks Data Engineer Professional, Azure DP-203 / DP-700, or
    AZ-305 certification.
  • dbt, Great Expectations or similar; Workday, Prism or Accounting
    Center exposure.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Data Engineer (Hands-On)
Lead Data Engineer (Hands-On)

Brickstech • Cary (NC)

Hybrid
USD 150,000 - 185,000
Lead Data Engineer – Azure Databricks
Lead Data Engineer – Azure Databricks

Solvimate Pvt. Ltd. • Cary (NC)

Hybrid
USD 140,000 - 145,000
Lead Data Engineer
Lead Data Engineer

K2 Integrity • New York (NY)

On-site
USD 140,000 - 210,000
Lead Data Engineer
Lead Data Engineer

Recru • Spring (TX)

On-site
USD 130,000 - 170,000
Senior Data Engineer with Databricks Exp. - 100% Remote
Senior Data Engineer with Databricks Exp. - 100% Remote

SDH Systems • United States

Remote
USD 140,000 - 190,000
Senior Azure Databricks Data Engineer
Senior Azure Databricks Data Engineer

EXL • New York (NY)

Hybrid
USD 140,000 - 190,000
Sr. Azure Data Engineer
Sr. Azure Data Engineer

Noblesoft Technologies Inc. • Greenfield (IN)

On-site
USD 120,000 - 180,000
Senior Data Engineer - Databricks
Senior Data Engineer - Databricks

DATAECONOMY • Raleigh (NC)

On-site
USD 120,000 - 160,000
Senior Data Engineer - Full Time Only - Remote
Senior Data Engineer - Full Time Only - Remote

GD Resources LLC • United States

Remote
USD 126,000 - 154,000
Remote work
Lead Databricks Engineer ( F2F interview | Parsippany, NJ | only W2 )
Lead Databricks Engineer ( F2F interview | Parsippany, NJ | only W2 )

COOLSOFT • Parsippany-Troy Hills (NJ)

On-site
USD 140,000 - 190,000