Lead Data Engineer

SIDRAM TECHNOLOGIES

Cary (NC)

Hybrid

USD 160,000 - 230,000

Full time

5 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Altimetrik is hiring a Lead Data Engineer for MetLife in Cary, NC. This hands-on leadership role requires deep expertise in building and scaling a lakehouse data platform on Azure Databricks, with production-grade PySpark and Scala code.

You will own end-to-end data architecture, AI-augmented ingestion, data quality, and governance. Expect mentoring, code reviews, and delivering reusable components for the client's data ecosystem.

Qualifications

  • Expert-level proficiency in Python, Scala, and PySpark with production-grade code experience.
  • Strong SQL, data modeling, and experience delivering lakehouse architectures.
  • Databricks Delta Lake, Unity Catalog, and Databricks Jobs/Workflows expertise.
  • Azure data stack including ADLS Gen2, Data Factory, and Event Hubs.
  • Experience designing and reviewing scalable data pipelines and CI/CD for data platforms.

Responsibilities

  • Own end-to-end data platform architecture for lakehouse (Bronze/Silver/Gold).
  • Design and implement CI/CD pipelines for Databricks and ADF using Azure DevOps.
  • Lead AI-augmented ingestion, canonical mapping, and data quality initiatives.
  • Build a semantic layer and knowledge-graph interfaces for data discovery.
  • Mentor engineers, conduct design reviews, and establish engineering patterns.

Skills

Python
Scala
PySpark
Databricks
Azure

Tools

Delta Lake
Unity Catalog
Azure Data Factory
ADLS Gen2

Job description

Lead Data Engineer

Altimetrik | Cary, NC | Hands-On | 12 18 Years of Experience

NOW HIRING: LEAD DATA ENGINEER
Client: MetLife
Company: Altimetrik
Location: Cary, NC Onsite / Hybrid
Employment: Full-Time
Experience: 15+ Years
Work Authorization: USC / H4EAD Only

No H1B Transfer

Hands-On Leadership Role
Lead Data Engineer

Altimetrik | Cary, NC | Hands-On | 12 18 Years of Experience

Role

Lead Data Engineer (Hands-On)

Location

Cary, NC (On-site / Hybrid)

Experience

12 18 Years Visa Independent candidates No H1 transfer

Employment

Full-Time with Altimetrik

Customer

MetLife

About The Engagement

Altimetrik is building a centralized, AI-first enterprise Data Hub for a global insurance and financial services client on Azure Databricks. The platform ingests 150+ inbound data feeds, distributes to 35+ downstream systems, and is organized as a medallion architecture (Bronze / Silver / Gold). AI is embedded in ingestion, canonical mapping, data quality, reconciliation, and business user access from day one not bolted on at the edges. This is a senior hands-on leadership role. You will own the end-to-end technical design of the data and AI layers, build the reference implementations your engineers work from, and ship production-grade Python, Scala, and PySpark code every week. If you have not written or reviewed production code in the past year, this is not the right fit.

What You Will Own
Data Platform Architecture & Engineering

Own the end-to-end lakehouse architecture: Bronze / Silver / Gold layer contracts, zone layout on ADLS Gen2, Delta Lake table design, partitioning, schema evolution, and retention strategy. Design and build metadata-driven, parameterized ingestion frameworks that onboard new data sources without bespoke pipeline code for every feed. Write canonical PySpark and Scala Spark transformation jobs that serve as the team reference; set coding and testing standards, review pull requests, and debug production incidents. Design for scale and cost: tune Spark clusters and pools, apply partition pruning and caching strategies, and set cost guardrails as data volumes grow. Build and automate CI/CD for Databricks and ADF pipelines in Azure DevOps using Databricks Asset Bundles and Terraform; maintain platform observability with Azure Monitor and Log Analytics. Design repeatable patterns for batch files, database extracts, CDC feeds, and streaming ingestion using Azure Event Hubs / Kafka and Spark Structured Streaming.

AI-Augmented Ingestion & Canonical Mapping

Design and build the AI-augmented metadata ingestion framework that auto-generates bridge documents, DML statements, canonical table definitions, and control metadata from source schemas. Build AI-assisted source-to-canonical attribute mapping: schema reasoning, data profiling, confidence scoring, and a human review / approval gate before any mapping reaches production. Generate file-level and record-level validation rules from historical data and metadata analysis, feeding a configurable, rules-engine-backed data quality framework.

AI-Driven Data Quality, Anomaly Detection & Testing

Build AI-assisted data quality that analyses patterns across Bronze, Silver, and Gold layers to propose DQ rules beyond predefined checks. Deliver anomaly detection covering outliers, data drift, schema drift, volume shifts, and reconciliation breaks with actionable alerting, not noise. Build AI-assisted automated reconciliation and test-data generation to feed the platform's automated testing framework. Produce synthetic, privacy-preserving datasets for lower environments using differential privacy, format-preserving masking, and referential-integrity-safe generation.

Semantic Layer & Conversational Data Access

Design and build an ontology-driven semantic layer and knowledge graph modelling relationships across finance data entities, powering data discovery, semantic integration, and AI/BI tooling. Build a GPT-powered conversational interface for natural-language querying of financial data: text-to-SQL or semantic-layer-mediated retrieval grounded in the knowledge graph, with row-level and column-level security enforced and every answer traceable to source.

Governance, Architecture Reviews & Team Leadership

Own end-to-end AI architecture decisions: model selection, RAG and retrieval design, prompt strategy, evaluation harnesses, guardrails, cost and latency budgets, and observability. Implement Unity Catalog for cataloguing, lineage, and fine-grained access control; define PII classification, masking, tokenization, and encryption standards across every layer. Take designs through Architecture Review Boards and AI governance forums, covering responsible AI, data residency, model approval, auditability, and human-in-the-loop controls. Mentor data engineers, run design reviews, and set the engineering patterns the team builds on without becoming a bottleneck. Produce documentation and reusable components good enough for the client's team to operate the platform independently at engagement end.

Must-have Skills & Experience
Programming & Data Engineering

Expert - level proficiency in Python, Scala, and PySpark, with a strong track record of designing and delivering production-ready, modular, and well-tested solutions; developing and troubleshooting Spark workloads; and optimizing large-scale batch and streaming data pipelines using Delta Lake and Spark technologies. Strong SQL and data modelling dimensional and normalised; schema design and data contract definition. Databricks expertise Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, Model Serving. Azure data stack ADLS Gen2 (zone design, ACLs, lifecycle), Azure Data Factory (parameterized / metadata-driven frameworks, error handling), Azure Event Hubs.

AI & Machine Learning

3+ years designing and shipping LLM-based systems in production: RAG pipelines, agentic / tool-calling workflows, structured output, chunking and embedding strategy, vector and hybrid retrieval, and prompt engineering. Evaluation discipline golden datasets, regression suites, accuracy and hallucination tracking, human-in-the-loop feedback loops; you measure AI quality, not assert it. Hands-on experience with LangChain, LlamaIndex, or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving). Metadata-driven thinking schema inference, data profiling, lineage, catalogs, and configuration-driven frameworks that onboard the next source without new code.

Architecture & Governance

12 18 years of total experience in data engineering, data platform delivery, or related disciplines. Proven delivery of a medallion / lakehouse architecture at enterprise scale not just familiarity with the concept. Azure security and governance Entra ID, managed identities, RBAC, POSIX ACLs on ADLS Gen2, Key Vault, private endpoints, and PII handling. CI/CD and infrastructure as code Azure DevOps, Terraform, Databricks Asset Bundles, and automated testing of data pipelines. Clear technical writing and the ability to present and defend a design to both engineers and non-technical stakeholders.

Strongly Preferred

Knowledge graphs and ontologies: RDF/SPARQL, property graphs (Neo4j), or graph modelling over a lakehouse. Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale, including access control and ambiguity handling. ML-based anomaly detection on time-series or transactional financial data. Financial services or insurance domain knowledge: finance close, general ledger, subledger, reconciliation, or actuarial data. LLMOps and MLOps: model versioning, prompt versioning, cost governance, and observability tooling. Databricks Data Engineer Professional, Azure DP-203 / DP-700, or AZ-305 certification. dbt, Great Expectations, or similar data-quality and transformation tooling. Workday, Prism, or Accounting Center exposure.

What Makes Someone Successful Here

You prototype in days, not sprints and the prototype is production-close enough to survive an architecture review. You know where AI genuinely helps and where a deterministic rule is the better engineering answer. On finance data, that judgement matters more than enthusiasm. You design for human review by default. Every AI-generated mapping, rule, and artefact lands in front of a reviewer with the reasoning attached. You are comfortable working with US-based client stakeholders and can explain a technical trade-off to a finance business owner without jargon. You leave behind documentation and patterns the client's own team can operate without you.

Hashtags

#Hiring #NowHiring #LeadDataEngineer #DataEngineer #SeniorDataEngineer #FullTime #USC #H4EAD #USCOnly #H4EADOnly #VisaIndependent #NoH1BTransfer #15PlusYears #15YearsExperience #15PlusYearsExperience #Python #Scala #PySpark #Databricks #Azure #ADLS #ADF #DeltaLake #Lakehouse #DataEngineering #DataArchitecture #GenerativeAI #GenAI #LLM #RAG #AgenticAI #LangChain #LangGraph #AzureOpenAI #UnityCatalog #DataQuality #MLOps #LLMOps #Terraform #AzureDevOps #MetLife #Altimetrik #CaryNC #NorthCarolina #FullTimeJobs #USITJobs

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Data Engineer (Hands-On)
Lead Data Engineer (Hands-On)

Cephas Consultancy Services Private Limited • Cary (NC)

Hybrid
USD 150,000 - 210,000
Lead Data Engineer (Hands-On)
Lead Data Engineer (Hands-On)

Brickstech • Cary (NC)

Hybrid
USD 150,000 - 185,000
Lead Data Engineer – Azure Databricks
Lead Data Engineer – Azure Databricks

Solvimate Pvt. Ltd. • Cary (NC)

Hybrid
USD 140,000 - 145,000
Lead Data Engineer
Lead Data Engineer

K2 Integrity • New York (NY)

On-site
USD 140,000 - 210,000
Sr. Azure Data Engineer
Sr. Azure Data Engineer

Noblesoft Technologies Inc. • Greenfield (IN)

On-site
USD 120,000 - 180,000
Lead Data Engineer
Lead Data Engineer

Recru • Spring (TX)

On-site
USD 130,000 - 170,000
Senior Data Engineer - Full Time Only - Remote
Senior Data Engineer - Full Time Only - Remote

GD Resources LLC • United States

Remote
USD 126,000 - 154,000
Remote work
Lead Data Engineer
Lead Data Engineer

Norfolk Southern Corp • Atlanta (GA)

Hybrid
USD 150,000 - 230,000
AI Data Engineer
AI Data Engineer

Howard Hughes Medical Institute • Kentucky

Hybrid
USD 120,000 - 180,000
Forward Deployed Data Engineer
Forward Deployed Data Engineer

Perform • Los Angeles (CA)

Hybrid
USD 170,000 - 210,000
Hybrid work model
Health benefits