AI Data Engineer

Howard Hughes Medical Institute

Kentucky

Hybrid

USD 120.000 - 180.000

Vollzeit

Vor 3 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Howard Hughes Medical Institute (HHMI) seeks an experienced AI Data Engineer to build and operate the data platforms powering AI across HHMI's knowledge and data fabric. You will implement pipelines, medallion architecture, and governance for AI-facing data assets, working under the Principal Knowledge & Data Architect.

This hybrid role requires hands-on production data engineering, Databricks/Spark expertise, and collaboration with the AI Developer and Operations teams.

Qualifikationen

  • Bachelor’s degree or equivalent; 4+ years of hands-on data engineering experience.
  • Production pipelines owned and operated in real environments.
  • Experience with AI/data foundation work is ideal.
  • Strong knowledge of Databricks platforms and cloud data tooling.

Aufgaben

  • Build the AI-facing data pipelines: ingestion, transformation, and governance-ready data.
  • Implement the medallion architecture with Delta Lake and scheduled materializations.
  • Own the workflow orchestration framework using Databricks Workflows and Delta Live Tables.
  • Implement governance patterns including Unity Catalog and access controls.
  • Build and operate retrieval-supporting infrastructure like embeddings and vector stores.
  • Partner with Operations Capabilities on source-system contracts and data interfaces.
  • Design and operate data quality, freshness, drift monitoring, and alerting.
  • Support AI Developer velocity by enabling data access for product teams.
  • Contribute to and evolve reusable reference patterns and templates.

Kenntnisse

Production data engineering
Databricks & Spark depth
Python
SQL
IaC with Terraform
AWS fundamentals

Ausbildung

Bachelor’s degree or equivalent

Tools

Databricks
Spark
Delta Lake
Delta Live Tables
Unity Catalog
Workflows
Airflow
Git
CI/CD
Terraform
AWS S3
KMS

Jobbeschreibung

Primary Work Address: 4000 Jones Bridge Road, Chevy Chase, MD, 20815

HHMI is focused on supporting and moving science forward in a variety of different ways ranging from conducting basic biomedical research, empowering educators, inspiring students, developing the next generation of scientists – even stretching into film and media production. Our Headquarters is in the greater Washington, DC metro area and is home to over 300 employees with expertise in investments, communications, digital production, biomedical sciences, and everything in between. The work housed here supports and augments the groundbreaking research conducted in HHMI labs across the nation. As HHMI scientists continue to push boundaries in laboratories and classrooms, you can be sure that your contributions while working here are making a difference.

The AI Accelerator exists to turn AI into daily reality across HHMI’s administrative and operational functions. This role builds the data-engineering foundation every AI application and knowledge layer at HHMI runs on.

The work is hands‑on. This person implements the pipelines, transformation patterns, and orchestration framework that turn HHMI’s institutional content into governed, AI‑ready data. They own the day‑to‑day execution of the data‑engineering side of the AI fabric — pipeline development, medallion curation, workflow‑orchestration framework, and governance implementation — under the design authority of the Principal Knowledge & Data Architect.

This role works closely with the Principal Knowledge & Data Architect (who owns the knowledge and retrieval layer’s design), the AI Developer (who consumes what this role builds), and Operations Capabilities’ Data Integration Engineer (who lands source‑system data at the interface). The AI Data Engineer is the seat that connects those layers into working data flow.

HHMI’s Principal K&D Architect designs the knowledge and data foundation for institutional AI, but the foundation only comes to life when someone builds and operates the pipelines that carry data through it. Without this role, our AI systems either wait on data that isn’t ready or consume content whose quality can’t be guaranteed. This role is what makes the K&D Architect’s design real day‑to‑day — and it’s the same role that keeps it working over years, as content evolves and models change.

This position works on a hybrid schedule, reporting in‑person three days a week to our headquarters in Chevy Chase, Maryland.

We encourage qualified candidates who are eligible to work in the United States.

Please note, we are not able to sponsor a visa for this position at this time.

What you will do
  • Build the AI‑facing data pipelines: ingestion from the landing zone Operations Capabilities delivers, transformation through raw → bronze → silver → gold, and serving of governed, AI‑ready content for downstream consumption.

  • Implement the medallion architecture: design patterns from the K&D Architect become working pipelines, tables, and materialization schedules. Delta Lake tables designed with partitioning, optimization, and evolution in mind.

  • Own the workflow‑orchestration framework: Databricks Workflows, Delta Live Tables, retry policies, alerting routes, run history, cost tags.

  • Implement governance patterns: Unity Catalog structure, sensitivity classification, access control, and audit for AI‑facing data assets. Design comes from the K&D Architect; day‑to‑day implementation lives here.

  • Build and operate retrieval‑supporting infrastructure: embedding pipelines, vector store maintenance, reindexing when models upgrade, retrieval evaluation frameworks.

  • Partner with Operations Capabilities on source‑system contracts: define what the AI Fabric consumes at the landing zone — schema, cadence, SLA, quality thresholds. Own the platform‑side of that contract.

  • Design and operate data quality: data‑quality checks, freshness monitoring, drift detection, and the alerting that surfaces issues before they hit AI users.

  • Support AI Developer velocity: when AI Developers deploy into product teams, this role is the data engineer they turn to when a use case needs specific data.

  • Contribute to and consume the reference‑pattern library: reusable pipeline patterns, code templates, and standards live in the shared platform layer. This role uses them, contributes new ones, and evolves them as we learn.

What We Are Looking For
  • Hands‑on production data engineering: at least four years designing, building, and operating production data pipelines. Not a Databricks‑course track record — real production experience where you owned the pipeline through breakage, iteration, and recovery.

  • Databricks and Spark depth: Delta Lake, medallion architecture, Delta Live Tables, Workflows, Databricks SQL, Unity Catalog. Comfortable at the layer where code meets platform.

  • Python and SQL fluency: PySpark, ETL patterns, and SQL that runs at scale. Version control (Git), CI/CD for data pipelines, and infrastructure‑as‑code (Terraform) as working tools.

  • AI‑adjacent data engineering: real experience building the data foundation for AI use cases — embedding pipelines, vector stores, chunking strategies, retrieval evaluation. Not required to be an ML researcher; required to have built the plumbing.

  • Workflow orchestration: Databricks Workflows, or Airflow in production. Retry semantics, dependency management, failure handling — as working discipline, not concepts.

  • Data quality and observability: Great Expectations, Databricks data‑quality monitors, or equivalent. Treats data quality as a first‑class engineering concern.

  • Governance discipline: works with Unity Catalog structures, understands sensitivity classification, and designs pipelines with access control and audit in mind from the first commit.

  • AWS foundations: IAM, S3, KMS at the level needed to work in a Databricks‑on‑AWS environment. Not required to be a cloud architect; required to be productive.

  • Communication: works productively with the K&D Architect on design, AI Developer on integration, and Operations Capabilities on contracts. Explains data‑engineering trade‑offs to non‑engineers.

  • Education and experience: bachelor’s degree or equivalent, plus at least four years of hands‑on data‑engineering experience with meaningful exposure to AI or knowledge‑management use cases.

Nice to Have
  • Prior experience with knowledge graphs (Neo4j or comparable), entity resolution, or semantic data models.

  • Experience with the modern data stack alongside Databricks‑native tooling.

  • Familiar

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Data Engineer
AI Data Engineer

HHMI • USA

Hybrid
USD 120.000 - 150.000
Hybrid work schedule
AI Data Engineer
AI Data Engineer

Howard Hughes Medical Institute • Bethesda (MD)

Hybrid
USD 129.000 - 161.000
AI Data Engineer
AI Data Engineer

Howard Hughes Medical Institute • Chevy Chase (MD)

Hybrid
USD 129.000 - 161.000
Hiring pay range: $128,816.80 - $161,0
Competitive pay
Exceptional health benefits
+3
AI Data Engineer
AI Data Engineer

Howard Hughes Medical Institute (HHMI) • Chevy Chase (MD), Northern (KY)

Hybrid
USD 129.000 - 161.000
Hybrid schedule
Competitive pay and benefits
Member of Data Staff (Analytics Engineer)
Member of Data Staff (Analytics Engineer)

United States Digital Space LLC • USA

Remote
USD 150.000 - 190.000
Data Engineer
Data Engineer

TDIndustries, Inc. • Dallas (TX)

Vor Ort
USD 108.000 - 132.000
AI Data Engineer
AI Data Engineer

Howard Hughes Medical Institute (HHMI) • USA

Hybrid
USD 129.000 - 161.000
Hybrid work schedule
Competitive compensation
Excellent health benefits
+1
Databricks Engineer
Databricks Engineer

CMT Services, Inc. • Adelphi (MD)

Vor Ort
USD 100.000 - 130.000
AI Data Engineer
AI Data Engineer

Delan Associates, Inc • USA

Vor Ort
USD 120.000 - 160.000
Sr. Data Engineer - AI
Sr. Data Engineer - AI

Dairy Farmers of America • Kansas City (KS)

Vor Ort
USD 120.000 - 180.000