Staff Data Engineer

Iterative Health

New York, Cambridge (NY, MA)

On-site

USD 200,000 - 325,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Iterative Health in New York is seeking an experienced Data Engineer to build and manage data infrastructure that supports predictive analytics in clinical trials. You'll work closely with ML and engineering teams to define data requirements, build pipelines, and ensure data quality. Ideal candidates will have over 10 years in data engineering, particularly with healthcare data systems, and be fluent in SQL and a modern programming language. The pay range is $200,000 - $325,000 USD.

Qualifications

  • 10+ years of experience in data engineering or related roles.
  • Significant experience with healthcare data (HL7, FHIR, etc.).
  • Strong SQL and at least one modern programming language.
  • Prior experience building systems from early stages.

Responsibilities

  • Own the data architecture and models.
  • Build data pipelines and transformations.
  • Ensure data quality and observability.
  • Work with ML teams to define data needs.

Skills

Data engineering experience
Healthcare data knowledge
SQL proficiency
Experience with AI and LLMs
Programming languages (Python, Java, etc.)

Tools

EHR systems
Data pipeline technologies
Cloud-native storage

Job description

About the Role

Accelerating clinical research is one of the defining challenges in healthcare. Promising therapies exist that patients can't access because the operational infrastructure to run clinical trials efficiently doesn't exist yet. We're building it. That means designing technology systems that bring order to a fragmented landscape of clinical data sources, automating the operational work that slows trials down, and turning real-world clinical data into a foundation for predictive intelligence.

We're building a uniquely valuable data asset: real-world patient and research data flowing across 80+ trial sites, spanning dozens of EHRs and clinical systems, focused on patient populations that are chronically underserved by existing clinical research infrastructure. Your job is to build the pipelines, data models, and AI infrastructure that make this asset real, from ingestion and normalization through to the systems that power predictions on top of it. You'll own data quality and observability as foundational engineering problems. You'll also have a direct hand in shaping how this data drives our AI strategy, what we model, what we predict, and what becomes possible.

This is an opportunity for someone who wants to be part of a small, fast-moving engineering team at a formative stage. You'll shape what gets built, how decisions get made, and what the team becomes.

Responsibilities
  • Own the data layer and architecture: the models, schemas, and infrastructure decisions that everything downstream depends on
  • Build and operate the pipelines and transformations that move data from ingestion through normalization, enrichment, and into the formats that support analytics, ML training, and production model serving
  • Own data quality and observability: build the systems that make data issues visible and correctable before they compound
  • Partner with ML and engineering teams to identify what's modelable, define training data requirements, and build the data foundations for new predictive capabilities
  • Define how clinical and operational data is governed across the system
  • Evaluate and select the tools and technologies that make up the data stack, with a clear point of view on build vs. buy
  • Help shape the engineering culture of a small, growing team: how technical decisions get made, how problems get debated, what rigor looks like in practice
Required Qualifications
  • 10+ years of experience in data engineering or related roles, with significant time spent building data systems
  • Experience with healthcare data strongly preferred (HL7, FHIR, claims, EHR extracts) or other complex, regulated data domains
  • Deep experience modeling and integrating data from multiple heterogeneous sources with inconsistent schemas and quality
  • Experience applying AI and LLMs to data engineering problems: extraction, normalization, classification, entity resolution
  • Strong understanding of how data infrastructure supports ML workflows from feature engineering to training data pipelines to model serving
  • Fluent in SQL and at least one modern programming language (Python, Java, Scala, Go), with experience across modern data infrastructure – distributed processing, streaming, cloud‑native storage, orchestration, and transformation frameworks
  • Have built data systems from early stages, making foundational decisions with incomplete information
  • Naturally raise the quality of the engineering around you through code review, design guidance, and honest technical conversation
Preferred Qualifications
  • Experience building data infrastructure that directly supports ML model training and evaluation
  • Familiarity with clinical trial operations, EDC systems, or life sciences data
  • SOC 2, HIPAA or similar compliance experience baked into engineering practice
  • A track record of building or improving data systems that others had given up on making reliable
Pay Range

$200,000 - $325,000 USD (New York)

Diversity & Inclusion

At Iterative Health, we’re actively working towards creating an environment that is representative of the diversity of patients our technology serves. We are focused on building an equitable and inclusive culture, and by extension, hiring process. If you require any accommodations to make the application process or interviewing experience more accessible to you, please contact CandidateAccommodations@iterative.health.

Equal Employment Opportunity

As set forth in Iterative Health’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Data Engineer
Staff Data Engineer

Iterativehealth • Town of Cambridge (NY)

On-site
USD 200,000 - 325,000
Staff Software Engineer
Staff Software Engineer

Iterativehealth • Town of Cambridge (NY)

On-site
USD 200,000 - 325,000
Staff Software Engineer
Staff Software Engineer

MD Ally • Cambridge (MA)

On-site
USD 200,000 - 325,000
Staff Data Scientist
Staff Data Scientist

Iterativehealth • Town of Cambridge (NY)

On-site
USD 200,000 - 325,000
Staff Data Scientist
Staff Data Scientist

MD Ally • Cambridge (MA)

On-site
USD 200,000 - 325,000
Equitable and inclusive culture
Accommodations for application process
Clinical Data Specialist
Clinical Data Specialist

Iterative Health • New York (NY)

On-site
USD 60,000 - 80,000
Comprehensive medical, dental, and vision coverage
Mental health and wellness support
Unlimited PTO and 12 company holidays
+1
Clinical Data Specialist
Clinical Data Specialist

Iterative Health • Southlake (TX)

Hybrid
USD 60,000 - 90,000
Hybrid work environment
Medical, dental, vision
401(k) with company match
+2
Clinical Data Specialist
Clinical Data Specialist

Iterative Scopes • Dallas (OR)

Hybrid
USD 55,000 - 75,000
Comprehensive medical, dental, and vision coverage
Unlimited PTO and 12 company holidays
401(k) program with company match
+1
Clinical Data Specialist
Clinical Data Specialist

Iterative Health • Cambridge (MA)

Hybrid
USD 70,000 - 100,000
Hybrid work arrangement with two in‑in
Medical, dental, and vision insurance
Wellness programs
+6
Director, Analytics
Director, Analytics

Iterative Health • Cambridge (MA)

On-site
USD 165,000 - 190,000