Lead Data Engineer - NBA

Humana

Boston (MA)

On-site

USD 140,000 - 200,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Humana seeks a Lead Data Engineer to own the NBA platform's data foundation on the Databricks Lakehouse, overseeing architecture, pipelines, and governance while remaining hands-on. You will lead a small team of engineers and contractors, building scalable batch and real-time pipelines from member, clinical, claims, behavioral, and socioeconomic datasets.

Responsibilities include designing Bronze/Silver/Gold data layers, implementing streaming ingestion with Kafka and Spark, and ensuring data

Qualifications

  • Bachelor's degree in computer science or related field.
  • 7 years of data engineering experience with at least 1-2 years in a lead engineer or technical leadership capacity.
  • Expert-level SQL and Python skills with significant experience developing and operating Spark-based data platforms.
  • Deep experience with Databricks, Delta Lake, and modern Lakehouse architectures.
  • Experience designing and operating Medallion Architecture (Bronze/Silver/Gold) implementations at enterprise scale.
  • Experience building and supporting both batch and streaming data pipelines in production environments.
  • Strong understanding of data quality engineering, validation frameworks, lineage, monitoring and operational support models.
  • Experience designing secure and compliant data platforms in regulated environments.
  • Demonstrated ability to lead a small engineering team while remaining an active hands-on contributor.
  • Strong communication skills and the ability to explain technical tradeoffs and architecture decisions to engineering, business, and compliance stakeholders.

Responsibilities

  • Own end-to-end delivery for the data engineering pod, including planning, execution, quality, and operational readiness.
  • Design and evolve the NBA Databricks Lakehouse, including Bronze, Silver, and Gold layer standards, data lineage, and governance patterns.
  • Design, build, and maintain scalable batch and real-time pipelines that process member, clinical, claims, behavioral, engagement, and socioeconomic datasets.
  • Lead the design and operation of Gold-layer feature tables and reusable data products that support model training, scoring, reinforcement learning, and decision intelligence workloads.
  • Design and implement Kafka- and Spark Structured Streaming-based ingestion and processing patterns that enable near real-time decision-making.
  • Establish platform-wide standards for data validation, monitoring, lineage, reconciliation, alerting, and operational visibility; treat data quality issues as production incidents.
  • Drive optimization of Spark workloads, Delta Lake storage patterns, partitioning strategies, and query performance to support enterprise-scale volumes efficiently.
  • Ensure compliance with PHI, HIPAA, auditability, retention, and governance requirements through Unity Catalog, lineage tooling, and secure data management practices.
  • Lead and mentor engineers and contractors; conduct code reviews, establish engineering standards, and drive adoption of best practices across the pod.
  • Partner with Decision Intelligence, Data Science, AI Engineering, and Platform Engineering teams to ensure reliable, governed, and performant data delivery across the NBA ecosystem.

Skills

SQL
Python
Databricks
Delta Lake
Medallion architecture
Spark
Data quality
Data governance
Leadership
Communication

Education

Bachelor's degree in computer science or related field

Tools

Databricks
Delta Live Tables
Unity Catalog
Kafka
Spark Structured Streaming

Job description

Become a part of our caring community

The Lead Data Engineer owns the NBA platform's data foundation on the Databricks Lakehouse. This role is responsible for the architecture, delivery, quality, governance, and operational reliability of the data products that power decision intelligence, machine learning, reinforcement learning, agentic AI, and real-time decisioning across the platform. You lead a small team of engineers and contractors while remaining deeply hands-on, building and evolving the pipelines, feature layers, streaming architectures, and data quality controls that turn raw healthcare data into trusted, production-ready assets consumed across the NBA ecosystem.

Key Responsibilities
  • Pod delivery --- Own end-to-end delivery for the data engineering pod, including planning, execution, quality, and operational readiness.
  • Lakehouse architecture --- Own the design and evolution of the NBA Databricks Lakehouse, including Bronze, Silver, and Gold layer standards, data lineage, and governance patterns.
  • Data pipeline engineering --- Design, build, and maintain scalable batch and real-time pipelines that process member, clinical, claims, behavioral, engagement, and socioeconomic datasets.
  • Feature platform ownership --- Lead the design and operation of Gold-layer feature tables and reusable data products that support model training, scoring, reinforcement learning, and decision intelligence workloads.
  • Streaming and event architecture --- Design and implement Kafka- and Spark Structured Streaming-based ingestion and processing patterns that enable near real-time decision-making.
  • Data quality and observability --- Establish platform-wide standards for data validation, monitoring, lineage, reconciliation, alerting, and operational visibility; treat data quality issues as production incidents.
  • Performance and optimization --- Drive optimization of Spark workloads, Delta Lake storage patterns, partitioning strategies, and query performance to support enterprise-scale volumes efficiently.
  • Data governance --- Ensure compliance with PHI, HIPAA, auditability, retention, and governance requirements through Unity Catalog, lineage tooling, and secure data management practices.
  • Team leadership --- Lead and mentor engineers and contractors; conduct code reviews, establish engineering standards, and drive adoption of best practices across the pod.
  • Cross-team coordination --- Partner with Decision Intelligence, Data Science, AI Engineering, and Platform Engineering teams to ensure reliable, governed, and performant data delivery throughout the NBA ecosystem.
Required Qualifications
  • Bachelor's degree in computer science or related field
  • 7 years of data engineering experience with at least 1--2 years in a lead engineer or technical leadership capacity.
  • Expert-level SQL and Python skills with significant experience developing and operating Spark-based data platforms.
  • Deep experience with Databricks, Delta Lake, and modern Lakehouse architectures.
  • Experience designing and operating Medallion Architecture (Bronze/Silver/Gold) implementations at enterprise scale.
  • Experience building and supporting both batch and streaming data pipelines in production environments.
  • Strong understanding of data quality engineering, validation frameworks, lineage, monitoring and operational support models.
  • Experience designing secure and compliant data platforms in regulated environments.
  • Demonstrated ability to lead a small engineering team while remaining an active hands-on contributor.
  • Strong communication skills and the ability to explain technical tradeoffs and architecture decisions to engineering, business, and compliance stakeholders.
Preferred Qualifications
  • Deep experience with Databricks Feature Store, Delta Live Tables, Unity Catalog, and Databricks Workflows.
  • Experience supporting machine learning, reinforcement learning, recommendation engines, or decision intelligence platforms.
  • Experience with Kafka, event streaming architectures, and real-time feature engineering pipelines.
  • Experience designing feature stores, feature-serving architectures, and model-data integration patterns.
  • Familiarity with Azure Data Factory, Azure Event Hubs, Azure Data Lake Storage, and broader Azure data platform services.
  • Experience implementing data observability and quality platforms such as Great Expectations, Monte Carlo, or equivalent solutions.
  • Experience integrating large external datasets including CMS, CDC, USDA, consumer, behavioral, or social determinants of health data sources.
  • Background in healthcare, insurance, or another regulated industry with PHI/HIPAA requirements.

We build with modern AI development tools (such as Claude and GitHub Copilot) and expect everyone on the team to use them to work faster and at higher quality.

Key Responsibilities
  • Architect, implement, and operate microservices that deliver:
    • Action and variant metadata
    • Context-aware policy and eligibility evaluation
    • Versioned, read-optimized APIs for high-performance runtime consumption
  • Guarantee that services are:
    • Highly available, with low latency
    • Horizontally scalable for increased demand
    • Backward compatible to support safe evolution and upgrades
  • Apply industry best practices for API design, schema evolution, service isolation, and secure integration.
Database & Schema Design
  • Design, deploy, and maintain resilient database schemas to support:
    • Comprehensive action and variant catalogs
    • Versioning, lifecycle management, and effective dating
    • Rule bindings and complex metadata relationships
  • Select and operate appropriate data stores (relational, document, key-value) tailored to workload and scalability requirements.
  • Implement and monitor:
    • Schema migration and backward compatibility strategies
    • Indexing and query optimization for performance
    • Data integrity, consistency, and reliability
    • Auditability and traceability for compliance and governance
Rules & Policy Engine Integration
  • Integrate and manage enterprise-grade rules engines to support:
    • Eligibility, constraints, and business policies
    • Suppression, cooldowns, exclusions, and other operational guardrails
    • Policy-driven allow/deny logic
  • Work with technologies such as Drools (DRL/DMN), IBM ODM, DMN-based services, OPA/Rego, or similar.
  • Ensure rule execution is deterministic, versioned, stateless, and free from unintended side effects.
AI-Assisted & Agentic Engineering
  • Utilize AI-powered and agentic tools to:
    • Generate and refactor database schemas and service logic
    • Streamline rule authoring, validation, and ongoing refactoring
    • Detect and address conflicting or redundant rules early in the development cycle
    • Automatically produce comprehensive test cases and explore edge scenarios
  • Apply AI responsibly to reduce manual effort while safeguarding clarity, correctness, and strong governance.
Testing, Reliability & Governance
  • Develop
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Data Engineer
Lead Data Engineer

K2 Intelligence, LLC • Northern (KY)

Hybrid
USD 120,000 - 180,000
Lead Data Engineer
Lead Data Engineer

K2 Integrity • New York (NY)

On-site
USD 140,000 - 210,000
Lead Decision Intelligence (AI) - NBA
Lead Decision Intelligence (AI) - NBA

Humana Inc • Lincoln (NE)

On-site
USD 120,000 - 180,000
Lead Decision Intelligence Engineer (AI) - NBA
Lead Decision Intelligence Engineer (AI) - NBA

Humana • United States

On-site
USD 140,000 - 210,000
Lead Decision Intelligence (AI) - NBA
Lead Decision Intelligence (AI) - NBA

Humana Inc • Dover (DE)

On-site
USD 180,000 - 240,000
Lead Data Engineer
Lead Data Engineer

Compunnel, Inc. • Northern (KY)

Hybrid
USD 120,000 - 170,000
Data Quality Engineer
Data Quality Engineer

Compunnel, Inc. • Chicago (IL), Northern (KY)

Hybrid
USD 110,000 - 165,000
Data Engineer
Data Engineer

Scorpion Therapeutics • Indianapolis (IN)

On-site
USD 120,000 - 180,000
Customer Solution Architect
Customer Solution Architect

Jobtailor • New York (NY)

On-site
USD 180,000 - 240,000
Databricks Engineer
Databricks Engineer

CMT Services, Inc. • Adelphi (MD)

On-site
USD 100,000 - 130,000