Staff Data Engineer- Data Lake

h1

New York (NY)

Hybrid

USD 170,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance options
Generous paid time off
Flexible work hours

Job summary

H1 in New York seeks a Staff Data Engineer to shape the architecture of their Data Lake. You will lead efforts in building robust data pipelines and ensure the quality of data used across the organization. This position offers a pathway to engineering management while remaining hands-on.

The ideal candidate has significant experience with distributed data platforms, a deep understanding of cloud-native systems, and a passion for mentoring teams. A comprehensive benefits package is included in this full-time role.

Qualifications

  • 8+ years of experience in data engineering or software engineering.
  • Demonstrated technical leadership and mentoring skills.
  • Strong proficiency in Python (PySpark) and advanced SQL expertise.

Responsibilities

  • Architect and build distributed ETL/ELT pipelines.
  • Lead the evolution of H1's Data Lake architecture.
  • Own and improve data quality across global sources.

Skills

Distributed data platforms
Data engineering
Python (PySpark)
Apache Spark
SQL performance tuning
Cloud-native environments
Containerization (Docker, Kubernetes)

Tools

AWS
Airflow
Argo

Job description

At H1, we believe access to the best healthcare information is a basic human right. Our mission is to provide a platform that can optimally inform every doctor interaction globally. This promotes health equity and builds needed trust in healthcare systems. To accomplish this, our teams harness the power of data and AI technology to unlock groundbreaking medical insights and convert those insights into actions that result in optimal patient outcomes and accelerate an equitable and inclusive drug development lifecycle. Visit h1.co to learn more about us.

Data Engineering is responsible for the development and delivery of our most important asset—our data. With thousands of data sources from around the world, the team ensures that data is accurate, normalized, and delivered at a velocity that keeps up with real-world changes. As we expand our markets and the scope of data we provide to our customers, our team must scale to meet that demand.

WHAT YOU'LL DO AT H1

As a Staff Data Engineer on the Data Lake team at H1, you will play a critical role in shaping the architecture, scalability, reliability, and long‑term direction of our core data platform. This role is designed for a highly technical engineer who is excited to grow into an Engineering Manager track while remaining deeply hands‑on technically.

The Data Lake is the foundation of H1’s platform, responsible for the validation, accuracy, standardization, and quality of the data powering every downstream product and team across the organization. You will help lead the evolution of this platform while supporting and mentoring a growing team of engineers.

You will:

  • Architect, build, and scale distributed ETL/ELT pipelines and large‑scale ingestion frameworks across structured and unstructured healthcare datasets.
  • Lead the evolution of H1’s Data Lake architecture with a focus on scalability, observability, reliability, and cost optimization.
  • Own and improve data quality, validation, normalization, and standardization workflows across thousands of global data sources.
  • Design and optimize batch and near real‑time data processing frameworks using cloud‑native distributed systems.
  • Optimize distributed compute and storage systems, including Spark workloads, query performance, partitioning strategies, and infrastructure efficiency.
  • Drive improvements in monitoring, governance, operational excellence, and production reliability across the platform.
  • Troubleshoot complex production data and infrastructure issues across distributed systems.
  • Partner closely with Product, Infrastructure, Security, Compliance, and downstream engineering teams to support scalable and secure data delivery.
  • Mentor engineers through technical leadership, architecture reviews, and engineering best practices.
  • Help define technical roadmap priorities and contribute to long‑term platform strategy and execution planning.
  • Support production operations, incident response, and platform health as part of overall ownership of the Data Lake ecosystem.
ABOUT YOU

You are a highly technical data engineer who thrives in lean, high‑ownership environments and enjoys solving complex distributed systems challenges. You are excited by the opportunity to influence technical direction, mentor engineers, and grow into broader engineering leadership responsibilities while remaining hands‑on.

  • You have deep experience designing and scaling distributed data platforms and large‑scale pipelines in cloud‑native environments.
  • You excel at building reliable, observable, and maintainable data systems supporting critical business and analytics workloads.
  • You have strong expertise in distributed processing, performance optimization, and modern data architecture patterns.
  • You are comfortable leading technical initiatives and influencing architecture decisions across teams.
  • You communicate effectively with both technical and non‑technical stakeholders.
  • You enjoy mentoring engineers and helping raise the engineering bar across teams.
  • You are energized by ownership, autonomy, and solving ambiguous technical challenges.
REQUIREMENTS
  • 8+ years of experience in data engineering, software engineering, or related fields with significant experience building and scaling distributed data platforms.
  • Demonstrated technical leadership experience with interest in or experience mentoring and leading engineers.
  • Strong proficiency in Python (PySpark), Java, Scala, or similar programming languages.
  • Advanced SQL expertise, including performance tuning and optimization across large datasets.
  • Deep experience with Apache Spark and cloud‑native big data platforms, preferably within AWS environments (EMR, Glue, S3, Athena, Redshift, or similar).
  • Experience designing and scaling modern cloud‑native data lake architectures and large‑scale ingestion frameworks.
  • Experience with orchestration and workflow management tools such as Argo, Airflow, or similar technologies.
  • Strong understanding of distributed storage systems, partitioning strategies, and file formats such as Parquet, Avro, and ORC.
  • Experience with Docker, Kubernetes, and modern containerization technologies.
  • Experience implementing monitoring, observability, and data quality frameworks within production environments.
  • Experience with large‑scale data cleaning, parsing, normalization, and validation workflows preferred.
  • Experience working with healthcare, life sciences, publication, or large‑scale entity‑resolution datasets preferred.
  • Exposure to ML/AI‑driven data enrichment, parsing, or validation workflows is a plus.
  • Experience using AI‑assisted coding tools (e.g., GitHub Copilot, Claude Code) to accelerate development while maintaining quality is encouraged.
COMPENSATION

This role pays $170,000 to $190,000 per year, based on experience, in addition to stock options.

Anticipated role close date: 8/1/2026

H1 OFFERS
  • Full suite of health insurance options, in addition to generous paid time off
  • Pre‑planned company‑wide wellness holidays
  • Retirement options
  • Health & charitable donation stipends
  • Impactful Business Resource Groups
  • Flexible work hours & the opportunity to work from anywhere
  • The opportunity to work with leading biotech and life sciences companies in an innovative industry with a mission to improve healthcare around the globe

H1 is proud to be an equal opportunity employer that celebrates diversity and is committed to creating an inclusive workplace with equal opportunity for all applicants and teammates. Our goal is to recruit the most talented people from a diverse candidate pool regardless of race, color, ancestry, national origin, religion, disability, sex (including pregnancy), age, gender, gender identity, sexual orientation, marital status, veteran status, or any other characteristic protected by law.

H1 is committed to working with and providing access and reasonable accommodation to applicants with mental and/or physical disabilities. If you require an accommodation, please reach out to your recruiter once you’ve begun the interview process. All requests for accommodations are treated discreetly and confidentially, as practical and permitted by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Data Engineer- RWE
Staff Data Engineer- RWE

H1 • New York (NY)

Hybrid
USD 170,000 - 190,000
Health insurance options
Generous paid time off
Pre-planned wellness holidays
+2
Staff Backend Engineer (RWE) - Distributed Systems & Data Platform
Staff Backend Engineer (RWE) - Distributed Systems & Data Platform

Transformcap • New York (NY)

On-site
USD 170,000 - 190,000
Health insurance
Stock options
Flexible work hours
Staff Data Engineer - Emerald
Staff Data Engineer - Emerald

h1 • New York (NY)

On-site
USD 170,000 - 190,000
Health insurance options
Flexible work hours
Generous paid time off
Sr. Product Manager, Digital Health Data Products
Sr. Product Manager, Digital Health Data Products

H1 • New York (NY)

On-site
USD 130,000 - 155,000
Full suite of health insurance options
Generous paid time off
Pre-planned company-wide wellness holidays
+3
Staff Data Engineer - Data & ML Platform
Staff Data Engineer - Data & ML Platform

Hinge Health • San Francisco (CA)

Hybrid
USD 197,000 - 297,000
Inclusive healthcare and benefits
401k retirement plan options with 2% match
Discounted company stock through ESPP
Senior Staff Data Engineer - Data & ML Platform
Senior Staff Data Engineer - Data & ML Platform

Hinge Health • San Francisco (CA)

Hybrid
USD 140,000 - 180,000
Comprehensive medical, dental, and vision coverage
401(k) plan with 2% company match
Learning and development stipends
+1
Enterprise Saas Implementation Specialist
Enterprise Saas Implementation Specialist

Transformcap • New York (NY)

On-site
USD 90,000 - 130,000
Health insurance and PTO
Retirement options
Wellness stipends
+2
Data Engineering Manager, Data & ML Platform
Data Engineering Manager, Data & ML Platform

Hinge Health • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Inclusive healthcare and benefits
2% company match on 401(k)
Learning and development stipends
Sr. Product Manager, Doctor & Research Site Profiles
Sr. Product Manager, Doctor & Research Site Profiles

h1 • New York (NY)

On-site
USD 130,000 - 160,000
Health insurance
Paid time off
Wellness holidays
+4
Sr. Product Engineer - API Developer
Sr. Product Engineer - API Developer

H1 • United States

On-site
USD 100,000 - 140,000
Full suite of health insurance options
Generous paid time off
Pre-planned company-wide wellness holidays
+4