AI Data Engineer: Delta Lake & Databricks Pipelines

Howard Hughes Medical Institute

Chevy Chase (MD)

Hybrid

USD 129,000 - 161,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Hiring pay range: $128,816.80 - $161,0
Competitive pay
Exceptional health benefits
Retirement plans
Time off
Wellness programs

Job summary

Howard Hughes Medical Institute (HHMI) invites applications for an AI Data Engineer to design, build, and operate governed, AI-ready data pipelines on HHMI’s Databricks-based platform. The role is based in Chevy Chase, MD and follows a hybrid model with three days per week in-person at HHMI offices.

You will own end-to-end pipelines, implement the medallion architecture, manage Delta Lake tables, and oversee data quality, governance, and retrieval frameworks to empower AI developers and

Qualifications

  • 4+ years of hands-on production data engineering experience designing, building, and operating production data pipelines.
  • Deep experience with Databricks and Spark including Delta Lake, medallion architecture, Delta Live Tables, Databricks Workflows, Databricks SQL, and Unity Catalog.
  • Proficiency in Python and SQL with PySpark and scalable ETL patterns.
  • Working knowledge of Git, CI/CD for data pipelines, and infrastructure-as-code with Terraform.
  • Experience building AI foundations: embedding pipelines, vector stores, retrieval evaluation frameworks.
  • Workflow orchestration experience with Databricks Workflows or Airflow in production.
  • Data quality and observability experience with Great Expectations or Databricks data-quality monitors.
  • Governance discipline with Unity Catalog structures, access control and audit from first commit.
  • AWS foundations for Databricks-on-AWS: IAM, S3, and KMS.
  • Strong communication skills to explain data-engineering trade-offs to non-engineers.
  • Bachelor’s degree or equivalent with exposure to AI or knowledge-management use cases.

Responsibilities

  • Build AI-facing data pipelines from ingestion to governed, AI-ready content for downstream use.
  • Implement the medallion architecture transforming patterns into working pipelines, tables, and materialization schedules.
  • Create and optimize Delta Lake tables with partitioning, optimization, and evolution.
  • Own workflow orchestration using Databricks Workflows and Delta Live Tables, including retries and cost tags.
  • Apply governance with Unity Catalog: classification, access control, and audit for AI data assets.
  • Develop retrieval-supporting infrastructure: embedding pipelines, vector stores, and retrieval evaluation frameworks.
  • Collaborate with Operations Capabilities to define source-system contracts (landing zone) including schema, cadence, and SLAs.
  • Design data quality and observability: checks, freshness, drift detection, and alerting.
  • Support AI Developer velocity by delivering data layers needed for specific use cases.
  • Contribute to and reuse the shared reference-pattern library with reusable pipelines and templates.

Education

Bachelor’s degree or equivalent

Tools

Databricks Workloads
Delta Live Tables
Delta Lake
Databricks SQL
Unity Catalog
Python
PySpark
SQL
Git
CI/CD
Terraform
Airflow
Great Expectations
AWS (IAM/S3/KMS)

Job description

Howard Hughes Medical Institute (HHMI) invites applications for an AI Data Engineer to design, build, and operate governed, AI-ready data pipelines on HHMI’s Databricks-based platform. The role is based in Chevy Chase, MD and follows a hybrid model with three days per week in-person at HHMI offices.

You will own end-to-end pipelines, implement the medallion architecture, manage Delta Lake tables, and oversee data quality, governance, and retrieval frameworks to empower AI developers and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Data Engineer — Build AI-Ready Data Pipelines
AI Data Engineer — Build AI-Ready Data Pipelines

Howard Hughes Medical Institute (HHMI) • Chevy Chase (MD), Northern (KY)

Hybrid
USD 129,000 - 161,000
Hybrid schedule
Competitive pay and benefits
AI Data Engineer — Build AI-Ready Data Pipelines
AI Data Engineer — Build AI-Ready Data Pipelines

Howard Hughes Medical Institute (HHMI) • United States

Hybrid
USD 129,000 - 161,000
Hybrid work schedule
Competitive compensation
Excellent health benefits
+1
AI Data Engineer: Build AI-Ready Data Pipelines (Hybrid)
AI Data Engineer: Build AI-Ready Data Pipelines (Hybrid)

Howard Hughes Medical Institute • Bethesda (MD)

Hybrid
USD 129,000 - 161,000
AI Data Engineer: Build AI-Ready Pipelines & Governance
AI Data Engineer: Build AI-Ready Pipelines & Governance

Howard Hughes Medical Institute • Kentucky

Hybrid
USD 120,000 - 180,000
AI Data Engineer — Build AI-Ready Data Pipelines
AI Data Engineer — Build AI-Ready Data Pipelines

HHMI • United States

Hybrid
USD 120,000 - 150,000
Hybrid work schedule
AI Data Engineer
AI Data Engineer

Howard Hughes Medical Institute • Chevy Chase (MD)

Hybrid
USD 129,000 - 161,000
Hiring pay range: $128,816.80 - $161,0
Competitive pay
Exceptional health benefits
+3
AI Data Engineer
AI Data Engineer

HHMI • United States

Hybrid
USD 120,000 - 150,000
Hybrid work schedule
Data Engineer: Lakehouse & Databricks Pipelines
Data Engineer: Lakehouse & Databricks Pipelines

InfoVision, Inc. • United States

On-site
USD 110,000 - 150,000
AI Data Engineer
AI Data Engineer

Howard Hughes Medical Institute (HHMI) • United States

Hybrid
USD 129,000 - 161,000
Hybrid work schedule
Competitive compensation
Excellent health benefits
+1
Senior AI Data Engineer: Databricks, Spark & ML Pipelines
Senior AI Data Engineer: Databricks, Spark & ML Pipelines

HCL Global Systems Inc • Princeton (NJ)

On-site
USD 150,000 - 190,000