The Howard Hughes Medical Institute (HHMI) is building an AI Accelerator, and this role helps power the foundation behind it. As an AI Data Engineer, you will design, build, and operate governed, AI-ready data pipelines that turn institutional content into data products for downstream retrieval and AI developer use. Work is performed on HHMI’s Databricks-based platform under the design authority of the Principal Knowledge & Data Architect.
This position is based in Chevy Chase, MD and works hybrid, with three days per week in-person at HHMI offices.
What you will do
- Build AI-facing data pipelines, including ingestion from Operations Capabilities landing zone inputs, transformation through raw bronze, silver, and gold layers, and serving of governed, AI-ready content for downstream consumption.
- Implement the medallion architecture by turning K&D Architect design patterns into working pipelines, tables, and materialization schedules.
- Build and operate Delta Lake tables with partitioning, optimization, and evolution in mind.
- Own workflow orchestration using Databricks Workflows and Delta Live Tables, including retry policies, alerting routes, run history, and cost tags.
- Implement governance patterns using Unity Catalog, including sensitivity classification, access control, and audit for AI-facing data assets.
- Build and operate retrieval-supporting infrastructure, including embedding pipelines, vector store maintenance, reindexing when models upgrade, and retrieval evaluation frameworks.
- Partner with Operations Capabilities to define source-system contracts for what the AI Fabric consumes at the landing zone, including schema, cadence, SLA, and quality thresholds.
- Partner with Operations Capabilities to own the platform-side of that contract.
- Design and operate data quality and observability: data-quality checks, freshness monitoring, drift detection, and alerting that surfaces issues before they reach AI users.
- Support AI Developer velocity by delivering the data layer AI Developers consume when use cases require specific data.
- Contribute to and consume the shared reference-pattern library, including reusable pipeline patterns, code templates, and platform standards.
Key requirements
- At least 4 years of hands-on production data engineering experience designing, building, and operating production data pipelines.
- Deep experience with Databricks and Spark, including Delta Lake, medallion architecture, Delta Live Tables, Databricks Workflows, Databricks SQL, and Unity Catalog.
- Python and SQL fluency, including PySpark and SQL that runs at scale, plus ETL patterns.
- Working knowledge of Git, CI/CD for data pipelines, and infrastructure-as-code with Terraform.
- AI-adjacent data engineering experience building foundations for AI use cases, including embedding pipelines, vector stores, chunking strategies, and retrieval evaluation.
- Workflow orchestration experience with Databricks Workflows or Airflow in production, including retry semantics, dependency management, and failure handling.
- Data quality and observability experience with Great Expectations, Databricks data-quality monitors, or equivalent.
- Governance discipline with Unity Catalog structures and sensitivity classification, designing for access control and audit from the first commit.
- AWS foundations at the level needed for Databricks-on-AWS work: IAM, S3, and KMS.
- Communication skills working with the K&D Architect, AI Developers, and Operations Capabilities, including explaining data-engineering trade-offs to non-engineers.
- Bachelor’s degree or equivalent, plus meaningful exposure to AI or knowledge-management use cases alongside the required data engineering experience.
Technologies
- Databricks Workflows, Delta Live Tables, Delta Lake, Databricks SQL, Unity Catalog
- Python, PySpark, SQL, Git
- CI/CD, Terraform
- Databricks data-quality monitors, Great Expectations
- Embedding pipelines, vector store maintenance, retrieval evaluation frameworks
- Databricks-on-AWS, IAM, S3, KMS
- Airflow
Compensation and benefits
- Hiring pay range: $128,816.80 - $161,021.00 (annual)
- Competitive pay
- Exceptional health benefits
- Retirement plans
- Time off
- Range of recognition and wellness programs
Reporting and employment details
- Reports to the Director of AI Enablement.
- HHMI is not able to sponsor a visa for this position at this time.
- HHMI uses E-Verify to confirm identity and employment eligibility for new hires.
Additional information
- The description outlines principal duties and responsibilities and is not necessarily exhaustive.
- Unless they begin with the word “may,” described essential duties and responsibilities are essential functions under the Americans with Disabilities Act.