Location: Cary, NC — On-site / HybridEmployment Type: Full-Time
Experience Required: 12–18 years
Eligibility: U.S. Citizens only
About the Engagement
This role supports the development of a centralized, AI-first enterprise Data Hub for a global insurance and financial services organization, built on Azure Databricks.
The platform ingests 150+ inbound data feeds, distributes data to 35+ downstream systems, and follows a modern medallion architecture (Bronze / Silver / Gold).
AI capabilities are embedded throughout the platform from day one, including ingestion, canonical mapping, data quality, reconciliation, anomaly detection, and business-user access.
This is a senior, hands-on technical leadership role. The Lead Data Engineer will own the end-to-end technical design of both the data and AI layers, establish reference implementations for the engineering team, and continue to write and review production-grade Python, Scala, and PySpark code every week.
Candidates who have not written or reviewed production code within the past year will not be considered.
Key Responsibilities
Data Platform Architecture & Engineering
- Own the enterprise lakehouse architecture, including:
- Bronze / Silver / Gold data contracts
- ADLS Gen2 zone design
- Delta Lake table architecture
- Partitioning strategies
- Schema evolution
- Retention policies
- Design and implement scalable, production-grade data engineering patterns for enterprise workloads.
Ingestion & Streaming Frameworks
- Build metadata-driven and parameterized ingestion frameworks supporting:
- Batch files
- Database extracts
- CDC feeds
- Real-time and streaming workloads
- Work extensively with:
- Azure Event Hubs
- Kafka
- Spark Structured Streaming
Hands-On Engineering Leadership
- Develop canonical PySpark and Scala Spark implementations.
- Establish engineering, coding, and automated testing standards.
- Conduct code and pull-request reviews.
- Troubleshoot production incidents and Spark workload failures.
- Optimize Spark clusters, workloads, and associated cloud costs.
- Define and enforce engineering guardrails for performance, scalability, and reliability.
CI/CD, Infrastructure & Observability
- Implement CI/CD pipelines for Databricks and Azure Data Factory using:
- Azure DevOps
- Databricks Asset Bundles
- Terraform
- Establish observability using:
- Azure Monitor
- Log Analytics
- Drive automated testing and repeatable deployment practices across environments.
AI-Augmented Data Engineering
- Design AI-assisted ingestion and canonical mapping solutions.
- Build capabilities for:
- Automated bridge-document generation
- DML generation
- Canonical table-definition generation
- AI-assisted source-to-canonical mapping
- Implement appropriate human-review and approval gates for AI-generated mappings and transformations.
AI-Driven Data Quality & Reconciliation
- Build AI/ML-based solutions for:
- Data drift detection
- Schema drift detection
- Volume anomaly detection
- Reconciliation failures
- Automated reconciliation
- Develop privacy-preserving synthetic test-data approaches.
Semantic Layer & Generative AI
- Help design the enterprise semantic layer and knowledge graph.
- Build GPT-powered conversational data-access capabilities, including:
- Text-to-SQL
- Semantic-layer retrieval
- RAG-based access patterns
- Enforce row-level and column-level security within AI-powered data-access solutions.
Governance & Technical Leadership
- Implement governance using Databricks Unity Catalog, including:
- Lineage
- Access control
- PII standards
- Data discovery and governance
- Participate in and lead:
- Architecture Review Boards
- AI governance forums
- Design reviews
- Mentor engineers and establish strong technical documentation standards.
- Present and defend architecture decisions to both technical and non-technical stakeholders.
Must-Have Skills & Experience
Core Data Engineering
- 12–18 years of total experience in data engineering, data architecture, or enterprise data-platform delivery.
- Proven experience delivering enterprise-scale medallion / lakehouse architectures.
- Expert-level hands-on experience with:
- Strong ability to build modular, well-tested, production-grade data solutions.
- Deep experience troubleshooting and tuning Spark workloads.
- Experience optimizing large-scale batch and streaming pipelines using Delta Lake.
SQL & Data Modeling
- Strong SQL expertise.
- Strong experience with:
- Dimensional data modeling
- Normalized data modeling
- Schema design
- Data contracts
Databricks
Strong hands-on experience with:
- Delta Lake
- Unity Catalog
- Databricks Jobs & Workflows
- Cluster and pool management
- Performance tuning
- Databricks Model Serving
Microsoft Azure Data Stack
Strong experience with:
- ADLS Gen2
- Zone architecture
- ACLs
- Lifecycle management
- Azure Data Factory
- Parameterized pipelines
- Metadata-driven ingestion frameworks
- Azure Event Hubs
Generative AI / LLM Engineering
- Minimum 3+ years of experience designing and deploying LLM-based systems in production.
- Strong hands-on experience with:
- RAG pipelines
- Agentic / tool-calling workflows
- Chunking strategies
- Embedding strategies
- Vector retrieval
- Hybrid retrieval
- Prompt engineering
- Experience with at least one of:
- LangChain
- LlamaIndex
- LangGraph
- Experience with at least one production AI stack:
- Azure OpenAI
- OpenAI
- Databricks Model Serving
LLM Evaluation & Quality
Experience establishing disciplined evaluation processes, including:
- Golden datasets
- Regression suites
- Accuracy measurement
- Hallucination tracking
- Human-in-the-loop feedback
Metadata & Governance
Strong experience building or working with metadata-driven frameworks involving:
- Schema inference
- Data profiling
- Lineage
- Data catalogs
Azure Security
Strong understanding of:
- Microsoft Entra ID
- Managed identities
- RBAC
- POSIX ACLs
- Azure Key Vault
- Private endpoints
- PII handling and data-security standards
DevOps & Infrastructure as Code
Hands-on experience with:
- Azure DevOps
- Terraform
- Databricks Asset Bundles
- Automated testing for data pipelines
Communication
- Strong technical writing skills.
- Ability to communicate complex architecture decisions clearly.
- Ability to present and defend technical designs to engineers, leadership, and non-technical stakeholders.
Strongly Preferred
Experience with any of the following is highly desirable:
- Knowledge graphs and ontologies
- RDF / SPARQL
- Neo4j
- Graph modeling over a lakehouse
- Enterprise-scale text-to-SQL solutions
- Semantic-layer-backed natural-language query platforms
- ML-based anomaly detection for:
- Time-series data
- Financial transactions
- Financial services or insurance domain experience, particularly:
- Finance close
- General Ledger
- Subledger
- Reconciliation
- Actuarial data
- LLMOps / MLOps, including:
- Model versioning
- Prompt versioning
- Cost governance
- Observability
- Certifications such as:
- Databricks Data Engineer Professional
- Microsoft Azure DP-203
- Microsoft Fabric DP-700
- AZ-305
- Experience with:
- dbt
- Great Expectations or similar data-quality platforms
- Workday
- Workday Prism
- Workday Accounting Center
Submission Requirements
Please include the following information:
- Updated resume
- Full Name
- Current Location
- Contact Number
- Email Address
- Work Authorization: U.S. Citizen
- LinkedIn Profile
- Availability to Start
- Interview Availability for the next 3 days
We are looking for a
true hands-on Lead Data Engineer who can combine enterprise architecture ownership with day-to-day production engineering. Candidates should be equally comfortable designing the target architecture, reviewing technical decisions, troubleshooting production workloads, and writing production code.