Staff Data Engineer (Emerald)

H1

York and North Yorkshire

Hybrid

GBP 90,000 - 130,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Flexible hours
Stock options
Health & life insurance
Unlimited PTO
Work from home

Job summary

H1 is seeking a Staff Data Engineer for the EMERALD team to shape architecture and scale for its healthcare entity resolution platform. You will lead a small team, stay hands-on, and own services powering automatching, identity mapping, grouping, deduplication, and enrichment across tens of millions of records.

You will collaborate with Product, AI/ML, Analytics, and Engineering to improve accuracy, reliability, and efficiency in a cloud-native environment using Spark, AWS, and modern infra

Qualifications

  • Strong hands-on data engineering background with distributed systems.
  • Experience building large-scale Spark/PySpark pipelines.

Responsibilities

  • Lead design, scalability, and technical direction of the EMERALD platform.
  • Own automatching, identity mapping, deduplication, and enrichment workflows.
  • Mentor engineers, review code, and set engineering best practices.
  • Collaborate with Product/AI/ML to improve matching precision and recall.

Skills

PySpark
Scala/Java
Apache Spark
AWS
Kubernetes
Terraform
Docker
ETL/ELT
Entity resolution
Distributed systems

Tools

Airflow

Job description

  • As a Staff Data Engineer on the Emerald team, you will play a critical role in shaping the architecture, scalability, and technical direction of H1’s healthcare entity resolution platform. EMERALD is responsible for linking large-scale external healthcare datasets, including PubMed, clinical trials, conferences, ct.gov, and web-collected data to H1’s canonical physician and organization profiles
  • This role sits at the intersection of distributed data engineering, entity matching, identity resolution, and large-scale healthcare data processing. You will lead a small team of engineers while remaining deeply hands‑on technically, owning the systems and pipelines powering automatching, grouping logic, identity mapping, deduplication, and enrichment workflows processing tens of millions of records
  • You will partner closely with Product, AI/ML, Analytics, and Engineering teams to improve platform accuracy, scalability, reliability, and operational efficiency across one of H1’s most critical data platforms
  • Lead the design, optimization, and scalability of distributed Spark/PySpark pipelines powering entity resolution and large-scale healthcare data processing
  • Own systems supporting automatching, identity mapping, grouping logic, deduplication, enrichment, and auto‑approval workflows across healthcare provider and organization datasets
  • Build and maintain scalable processing frameworks for PubMed, clinical trial, ct.gov, conference, and other healthcare data sources
  • Drive infrastructure optimization initiatives focused on improving throughput, runtime, observability, and cloud compute cost efficiency
  • Partner closely with AI/ML teams to integrate matching and resolution models into EMERALD and improve matching precision and recall
  • Lead complex technical initiatives from architecture and design through deployment, monitoring, and long‑term production support
  • Serve as a technical leader and mentor across the team through code reviews, technical guidance, and engineering best practices
  • Collaborate directly with Product and business stakeholders to align technical solutions with operational and customer needs
  • Support production operations, incident response, troubleshooting, and ongoing platform reliability
Benefits
  • Flexible work hours
  • Commuter benefits
  • Stock options
  • Computer setup
  • Work from home opportunities
  • Health & life insurance
  • Retirement options
  • Unlimited PTO
  • Flex Give & Flex Spend
  • Impactful BRGs

Experience building scalable ETL/ELT frameworks across both batch and streaming architecturesExtensive experience with Apache Spark and AWS-based big data technologies including EMR, S3, and distributed compute environmentsExperience optimizing large-scale Spark workloads for performance, scalability, and infrastructure cost efficiencyExperience with containerization and infrastructure technologies such as Docker, Kubernetes, and TerraformExperience working with relational or distributed databases such as PostgreSQL or RedshiftYou bring strong hands‑on engineering expertise across distributed computing, large‑scale data processing, and infrastructure optimization while also helping guide technical direction and mentor engineers across the organizationStrong grasp of software engineering fundamentals including distributed systems, data structures, concurrency, and system designExperience working with healthcare, life sciences, Real World Evidence (RWE), or large‑scale healthcare datasets is strongly preferredStrong coding experience in Python (PySpark), Scala, Java, or equivalent languages used for distributed processing systemsDeep expertise with distributed data processing frameworks such as Apache Spark and Hadoop, particularly within AWS environmentsStrong communication and collaboration skills across both technical and non‑technical stakeholdersExperience improving performance, scalability, observability, and infrastructure efficiency within distributed systemsExperience with streaming and event‑driven architectures using technologies such as Kafka or Spark StreamingExperience with entity resolution, identity mapping, automatching, deduplication, or large‑scale matching systems is strongly preferredProven ability to operate effectively within highly scalable, production‑grade distributed systemsFamiliarity with modern development and infrastructure tooling including Git, CI/CD pipelines, Docker, Kubernetes, Terraform, Argo, Hudi, and JIRA8+ years of experience building and maintaining large‑scale distributed data systems and pipelinesExperience with streaming technologies such as Kafka, Spark Streaming, or KSQLExperience performing root cause analysis across large‑scale distributed systems and complex data pipelinesStrong proficiency in Python (PySpark), Scala, Java, or other modern programming languages used for large‑scale distributed processingDemonstrated technical leadership experience mentoring engineers and driving complex technical initiativesExperience with orchestration and lakehouse technologies such as Argo and Hudi or comparable platformsYou are an experienced data engineer with deep expertise building and optimizing distributed data systems in cloud‑native environments. You thrive solving complex scalability and performance challenges across high‑volume data processing systems and enjoy operating in highly technical, fast‑paced engineering environmentsAbility to write clean, maintainable, modular, and production‑grade codeStrong understanding of distributed file formats including Apache Parquet and Apache AVRO

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Data Engineer — Spark & Healthcare Data (Remote)
Staff Data Engineer — Spark & Healthcare Data (Remote)

H1 • York and North Yorkshire

Hybrid
GBP 90,000 - 130,000
Flexible hours
Stock options
Health & life insurance
+2
Senior Data Engineer
Senior Data Engineer

Our Future Health • Greater London

On-site
GBP 75,000 - 110,000
Principal Data Engineer
Principal Data Engineer

Medidata Solutions • Greater London

Hybrid
GBP 110,000 - 150,000
Medical insurance
Dental insurance
Life insurance
+3
Principal Data Engineer
Principal Data Engineer

MediData • Greater London

Hybrid
GBP 120,000 - 180,000
Medical insurance
Dental insurance
Life insurance
+3
Staff Data Engineer
Staff Data Engineer

twentysix • Washington

On-site
GBP 55,000 - 80,000
Senior Data Engineer
Senior Data Engineer

NextGenEnergyJobs • Greater London

Hybrid
GBP 65,000 - 90,000
Data Engineer
Data Engineer

Athsai • Greater London

On-site
GBP 50,000 - 70,000
Data Engineer (Recruiting Analytics)
Data Engineer (Recruiting Analytics)

Anthropic • York and North Yorkshire

Hybrid
GBP 90,000 - 120,000
Health insurance
Fertility benefits
Parental Leave 22 weeks
+4
Staff Software Engineer - Data Platforms
Staff Software Engineer - Data Platforms

Our Future Health UK • Greater London

On-site
GBP 120,000 - 180,000
Senior Data Engineer
Senior Data Engineer

Luxoft • Greater London

Hybrid
GBP 70,000 - 110,000