Lead Data Engineer

Nazztec group

Cary (NC)

On-site

USD 180,000 - 320,000

Full time

8 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

MetLife Cary, NC seeks a Lead Data Engineer – Hands-On to own the end-to-end data and AI architecture. You will mentor engineers, deliver production Python/Scala/PySpark code, and drive the AI-driven data hub across 150+ feeds and 35+ downstream systems.

You will implement Bronze/Silver/Gold contracts, CI/CD with Azure DevOps, and IaC with Terraform while optimizing Spark workloads and ensuring data governance and security.

Qualifications

  • Total 12–18 years in Data Engineering / Data Platform delivery.
  • Hands-on with production-grade Python, Scala and PySpark in modern lakehouse environments.
  • Experience building AI-first data hubs and enterprise-scale ingestion pipelines.
  • Proven ability to mentor engineers and write/review production code within the last year.
  • Strong security, governance and metadata capabilities across Unity Catalog and RBAC.

Responsibilities

  • Design and implement enterprise Lakehouse architecture with Bronze/Silver/Gold.
  • Own end-to-end data and AI layer design and production code delivery.
  • Lead CI/CD for Databricks and IaC with Terraform and Azure DevOps.
  • Mentor teams, conduct code reviews, optimize Spark workloads and pipelines.
  • Deliver production-grade AI/LLM solutions and governance frameworks.

Skills

Azure Databricks
Python
Scala
PySpark
Spark
Data modeling
Data governance
Mentoring
Architecture design

Tools

Unity Catalog
Databricks Job/Workflows
Terraform
Azure DevOps
ADLS Gen2
LangChain
LangGraph
Databricks Model Serving

Job description

ROLE: Lead Data Engineer – Hands-On

Client: MetLife

Location: Cary, NC — On-site / Hybrid

Experience: 12–18 Years

Employment Type: Full-Time

The engagement involves building a centralized, AI-first enterprise Data Hub for a global insurance and financial services organization using Azure Databricks.

The platform ingests 150+ inbound data feeds and distributes data to 35+ downstream systems, leveraging a Medallion Architecture — Bronze / Silver / Gold.

AI is embedded across ingestion, canonical mapping, data quality, reconciliation, and business-user access.

This is a senior hands-on technical leadership role. The selected candidate will own the end-to-end technical design of the data and AI layers, build reference implementations, mentor engineering teams, and continue writing/reviewing production-grade Python, Scala, and PySpark code every week.

Candidates who have NOT written or reviewed production code within the past year will NOT be considered.

Data Platform Architecture & Engineering
  • Design and implement enterprise-scale Lakehouse architecture.
  • Define Bronze / Silver / Gold data contracts.
  • Build scalable and reusable data platform components.
Metadata-Driven Ingestion
  • Design parameterized ingestion frameworks for:
  • Batch files
  • CDC feeds
  • Build schema inference, data profiling, metadata and lineage capabilities.
  • Develop production-grade Python, Scala and PySpark solutions.
  • Establish coding and testing standards.
  • Conduct PR/code reviews.
  • Optimize Spark workloads, clusters and pipelines.
  • Implement performance and cost optimization strategies.
  • Unity Catalog
  • Databricks Jobs & Workflows
  • Cluster and pool management
  • Model Serving
  • ADLS Gen2
CI/CD & Infrastructure as Code
  • Implement CI/CD for Databricks and ADF using Azure DevOps.
  • Work with Databricks Asset Bundles.
  • Implement Infrastructure as Code using Terraform.
  • Automate data pipeline testing and deployments.
  • Establish observability and operational monitoring.

A major component of this role is designing and delivering production-grade AI/LLM solutions.

You will work on:
  • AI-augmented ingestion.
  • AI-assisted canonical mapping.
  • Auto-generated bridge documentation.
  • DML and canonical table generation.
  • Source-to-canonical mapping with human review gates.
  • Data drift and schema drift detection.
  • Volume anomaly detection.
  • Automated reconciliation.
  • Synthetic privacy-preserving test data generation.
  • Knowledge graphs.
  • GPT-powered conversational interfaces.
  • Text-to-SQL.
  • Row-level and column-level security.

Candidates must have 3+ years of hands-on experience designing and shipping LLM-based systems in production, including:

  • Agentic AI workflows
  • Tool-calling workflows
  • Chunking strategies
  • Embedding strategies
  • Vector retrieval
  • Hybrid retrieval
  • LLM evaluation
  • Regression test suites
  • Accuracy tracking
  • Hallucination monitoring
  • Human-in-the-loop feedback
Hands-on experience with at least one of:
  • LangChain
  • LangGraph
Provider/platform experience:
  • OpenAI
  • Databricks Model Serving
Strong enterprise security and governance experience is required across:
  • Unity Catalog
  • Access control
  • PII standards
  • Managed Identities
  • RBAC
  • POSIX ACLs
  • Data privacy and protection
The candidate will also participate in:
  • AI Governance Forums

12–18 years of total experience in Data Engineering / Data Platform delivery

Data contracts & schema design

Unity Catalog

Kafka / Spark Structured Streaming

Spark performance tuning

Enterprise Lakehouse / Medallion Architecture

Azure security & governance

3+ years production LLM/GenAI engineering

RAG / Agentic AI / Tool Calling

LangChain / LlamaIndex / LangGraph

Azure OpenAI / OpenAI / Databricks Model Serving

  • RDF / SPARQL
  • Neo4j
  • Graph modelling over Lakehouse
  • Text-to-SQL
  • ML-based anomaly detection
  • Finance Close
  • Subledger
  • LLMOps / MLOps
  • LLM cost governance
  • AI observability
  • dbt
  • Workday
  • Workday Prism
  • Accounting Center

We are looking for a hands-on Lead Data Engineer / Data Architect / Principal Data Engineer who can operate at both architecture and engineering levels.

You should be comfortable designing the architecture, making technical decisions, writing production code, reviewing code, troubleshooting Spark workloads, mentoring engineers, and presenting architecture decisions to senior technical and business stakeholders.

This is NOT a pure Data Architect, Solution Architect, Engineering Manager, or AI Strategy role.

Recent hands-on production coding experience is mandatory.

#Nazztec #NazztecHiring #JoinNazztec #Hiring #UrgentHiring #LeadDataEngineer #DataEngineer #SeniorDataEngineer #PrincipalDataEngineer #DataArchitect #DataEngineering #DataPlatform #DataPlatformEngineer #BigData #BigDataEngineering #Azure #MicrosoftAzure #AzureDatabricks #Databricks #DeltaLake #UnityCatalog #AzureDataFactory #ADF #ADLS #ADLSGen2 #AzureEventHubs #Kafka #Spark #ApacheSpark #PySpark #Scala #Python #SQL #DataLake #Lakehouse #LakehouseArchitecture #MedallionArchitecture #BronzeSilverGold #DataArchitecture #DataModeling #DataEngineeringJobs #AzureJobs #DatabricksJobs #AI #ArtificialIntelligence #GenerativeAI #GenAI #LLM #LLMEngineering #LLMEngineer #LLMJobs #RAG #RAGPipeline #RetrievalAugmentedGeneration #AgenticAI #AgenticWorkflows #ToolCalling #LangChain #LlamaIndex #LangGraph #AzureOpenAI #OpenAI #DatabricksModelServing #PromptEngineering #VectorDatabase #VectorSearch #HybridSearch #Embeddings #SemanticSearch #TextToSQL #SemanticLayer #KnowledgeGraph #Neo4j #RDF #SPARQL #MLOps #LLMOps #AIEngineering #AIEngineer #MachineLearning #AnomalyDetection #DataQuality #DataGovernance #DataSecurity #DataLineage #DataCatalog #Metadata #MetadataDriven #DataContracts #SchemaEvolution #DataProfiling #DataObservability #DataOps #DevOps #AzureDevOps #Terraform #InfrastructureAsCode #IaC #CICD #ContinuousIntegration #ContinuousDeployment #SparkTuning #PerformanceTuning #CloudDataEngineering #CloudArchitecture #EnterpriseData #EnterpriseArchitecture #FinancialServices #Insurance #InsuranceTechnology #FinTech #MetLife #CaryNC #NorthCarolinaJobs #USJobs #USAJobs #USCitizen #USCitizenJobs #TechJobs #ITJobs #TechnologyJobs #DataScience #CloudJobs #DataProfessionals #TechHiring #ITHiring #DataHiring #AIJobs #RemoteJobs #HybridJobs #OnsiteJobs

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Data Engineer
Lead Data Engineer

SIDRAM TECHNOLOGIES • Cary (NC)

Hybrid
USD 160,000 - 230,000
Lead Data Engineer (Hands-On)
Lead Data Engineer (Hands-On)

Brickstech • Cary (NC)

Hybrid
USD 150,000 - 185,000
Lead Data Engineer
Lead Data Engineer

Accrescent Group • Cary (NC)

On-site
USD 140,000 - 180,000
Lead Data Engineer – Azure Databricks
Lead Data Engineer – Azure Databricks

Solvimate Pvt. Ltd. • Cary (NC)

Hybrid
USD 140,000 - 145,000
Lead Databricks ML & AI Ops Engineer
Lead Databricks ML & AI Ops Engineer

KData Inc. • United States

Remote
USD 180,000 - 240,000
Lead Data Engineer - NBA
Lead Data Engineer - NBA

Humana • Boston (MA)

On-site
USD 140,000 - 200,000
Data Architect / Lead Data Engineer
Data Architect / Lead Data Engineer

Mind Ware Inc • United States

Remote
USD 120,000 - 180,000
Senior Databricks Data Engineer
Senior Databricks Data Engineer

CRED • College Station (TX)

On-site
USD 120,000 - 160,000
Incentive Program
Accrued Time Off
401(k) with match
+4
Lead Data Engineer
Lead Data Engineer

Norfolk Southern Corp • Atlanta (GA)

Hybrid
USD 150,000 - 230,000
Senior Data Engineer - Agentic AI, Automation, and Data Platforms
Senior Data Engineer - Agentic AI, Automation, and Data Platforms

General Motors • Warren (MI)

Hybrid
USD 140,000 - 170,000