Staff Data Engineer

TAU Ventures

Massachusetts

On-site

USD 200,000 - 325,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Iterative Health, a healthcare technology and services company, builds data pipelines and AI infrastructure to accelerate clinical research and improve patient outcomes. This role sits on a small, fast-moving engineering team at a formative stage, shaping data models, governance, and the ML stack across 80+ trial sites.

Based in Cambridge, MA and New York City, the company emphasizes impact and collaboration in a growing environment.

Qualifications

  • 10+ years of experience in data engineering or related roles.
  • Experience with healthcare data (HL7, FHIR, claims, or EHR extracts) preferred.
  • Experience modeling and integrating data from multiple heterogeneous sources.
  • Experience applying AI and LLMs to data engineering problems.
  • Strong ML workflow understanding from feature engineering to training data pipelines to model serving.
  • Fluent in SQL and at least one modern programming language (Python, Java, Scala, Go) with experience across modern data infrastructure.
  • Built data systems from early stages, making foundational decisions with incomplete information.
  • Promotes high engineering quality via code reviews and design guidance.

Responsibilities

  • Own the data layer and architecture: the models, schemas, and infrastructure decisions that everything downstream depends on.
  • Build and operate the pipelines and transformations that move data from ingestion through normalization, enrichment, and into the formats that support analytics, ML training, and production model serving.
  • Own data quality and observability: build the systems that make data issues visible and correctable before they compound.
  • Partner with ML and engineering teams to identify what's modelable, define training data requirements, and build the data foundations for new predictive capabilities.
  • Define how clinical and operational data is governed across the system.
  • Evaluate and select the tools and technologies that make up the data stack, with a clear point of view on build vs buy.
  • Help shape the engineering culture of a small, growing team: how technical decisions get made, how problems get debated, what rigor looks like in practice.

Skills

Data engineering
Healthcare data
Data modeling
AI/LLMs in data engineering
ML workflows
SQL
Python

Tools

HL7
FHIR

Job description

Iterative Health is a healthcare technology and services company powering the acceleration of clinical research to transform patient outcomes.

We built a leading performance‑driven network of 100+ sites across the US, Europe, India, and Australia, conducting research directly in the communities where care is delivered across gastrointestinal, hepatology, obesity, and cardiology. By combining deep clinical trial expertise with cutting‑edge AI, we connect sponsors' scientific ambitions with high‑performing research teams that expedite and expand access to novel therapeutics for patients in need. Today, Iterative Health is headquartered in Cambridge, Massachusetts, and New York City with 250+ employees world‑wide.

About the Role

Accelerating clinical research is one of the defining challenges in healthcare. Promising therapies exist that patients can't access because the operational infrastructure to run clinical trials efficiently doesn't exist yet. We're building it. That means designing technology systems that bring order to a fragmented landscape of clinical data sources, automating the operational work that slows trials down, and turning real-world clinical data into a foundation for predictive intelligence.

We're building a uniquely valuable data asset: real‑world patient and research data flowing across 80+ trial sites, spanning dozens of EHRs and clinical systems, focused on patient populations that are chronically underserved by existing clinical research infrastructure. Your job is to build the pipelines, data models, and AI infrastructure that make this asset real, from ingestion and normalization through to the systems that power predictions on top of it. You'll own data quality and observability as foundational engineering problems. You'll also have a direct hand in shaping how this data drives our AI strategy, what we model, what we predict, and what becomes possible.

This is an opportunity for someone who wants to be part of a small, fast‑moving engineering team at a formative stage. You'll shape what gets built, how decisions get made, and what the team becomes.

Responsibilities
  • Own the data layer and architecture: the models, schemas, and infrastructure decisions that everything downstream depends on
  • Build and operate the pipelines and transformations that move data from ingestion through normalization, enrichment, and into the formats that support analytics, ML training, and production model serving
  • Own data quality and observability: build the systems that make data issues visible and correctable before they compound
  • Partner with ML and engineering teams to identify what's modelable, define training data requirements, and build the data foundations for new predictive capabilities
  • Define how clinical and operational data is governed across the system
  • Evaluate and select the tools and technologies that make up the data stack, with a clear point of view on build vs. buy
  • Help shape the engineering culture of a small, growing team: how technical decisions get made, how problems get debated, what rigor looks like in practice
What We’re Looking For

Required Qualifications

  • 10+ years of experience in data engineering or related roles, with significant time spent building data systems
  • Experience with healthcare data strongly preferred (HL7, FHIR, claims, EHR extracts) or other complex, regulated data domains
  • Deep experience modeling and integrating data from multiple heterogeneous sources with inconsistent schemas and quality
  • Experience applying AI and LLMs to data engineering problems: extraction, normalization, classification, entity resolution
  • Strong understanding of how data infrastructure supports ML workflows from feature engineering to training data pipelines to model serving
  • Fluent in SQL and at least one modern programming language (Python, Java, Scala, Go), with experience across modern data infrastructure - distributed processing, streaming, cloud‑native storage, orchestration, and transformation frameworks
  • Have built data systems from early stages, making foundational decisions with incomplete information
  • Naturally raise the quality of the engineering around you through code review, design guidance, and honest technical conversation

Preferred Qualifications

  • Experience building data infrastructure that directly supports ML model training and evaluation
  • Familiarity with clinical trial operations, EDC systems, or life sciences data
  • SOC 2, HIPAA or similar compliance experience baked into engineering practice
  • A track record of building or improving data systems that others had given up on making reliable

New York pay range

$200,000 — $325,000 USD

At Iterative Health, we’re actively working towards creating an environment that is representative of the diversity of patients our technology serves. We are focused on building an equitable and inclusive culture, and by extension, hiring process. If you require any accommodations to make the application process or interviewing experience more accessible to you, please contact CandidateAccommodations@iterative.health.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Data Engineer
Staff Data Engineer

Iterative Health • New York (NY), Cambridge (MA)

On-site
USD 200,000 - 325,000
Applied AI Engineer
Applied AI Engineer

TAU Ventures • Cambridge (MA)

Hybrid
USD 170,000 - 240,000
Hybrid work
Medical coverage
Wellness programs
+1
Data Manager
Data Manager

Iterative Health • New York (NY), Town of Texas (WI)

Hybrid
USD 125,000 - 160,000
Hybrid work
Medical, dental, vision
Mental health support
+4
Data Manager
Data Manager

Iterative-Health • Southlake (TX)

Hybrid
USD 125,000 - 160,000
Hybrid work two days per week in NYC,南
Hybrid work two days per week in South
Hybrid work two days per week inBoston
Data Manager
Data Manager

Iterative-Health • New York (NY)

Hybrid
USD 125,000 - 160,000
Hybrid work two days in NYC/Southlake?
Medical, dental, vision coverage
Mental health support
+4
Data Manager
Data Manager

Iterative Health • Southlake (TX)

Hybrid
USD 125,000 - 160,000
Hybrid work environment
Medical, dental, vision coverage
401(k) with match
Clinical Data Specialist
Clinical Data Specialist

TAU Ventures • Dallas (OR)

On-site
USD 65,000 - 90,000
Hybrid work environment with in-office
Medical, dental, and vision coverage
401(k) with company match
+4
Data Manager
Data Manager

TAU Ventures • Southlake (TX)

Hybrid
USD 125,000 - 160,000
Hybrid work two days/week
Medical, dental, and vision
Mental health support
+2
Technical Services Engineer – Site Infrastructure & Enablement
Technical Services Engineer – Site Infrastructure & Enablement

Iterative Health • Southlake (TX)

On-site
USD 85,000 - 115,000
Hybrid work environment
Medical, dental, vision coverage
401(k) with company match
+2
Clinical Data Specialist
Clinical Data Specialist

Iterative Health • New York (NY)

On-site
USD 60,000 - 80,000
Comprehensive medical, dental, and vision coverage
Mental health and wellness support
Unlimited PTO and 12 company holidays
+1