An application made for this job — a tailored resume and cover letter that speak straight to the posting.
H1 is seeking a Staff Data Engineer for the EMERALD team to shape architecture and scale for its healthcare entity resolution platform. You will lead a small team, stay hands-on, and own services powering automatching, identity mapping, grouping, deduplication, and enrichment across tens of millions of records.
You will collaborate with Product, AI/ML, Analytics, and Engineering to improve accuracy, reliability, and efficiency in a cloud-native environment using Spark, AWS, and modern infra
Experience building scalable ETL/ELT frameworks across both batch and streaming architecturesExtensive experience with Apache Spark and AWS-based big data technologies including EMR, S3, and distributed compute environmentsExperience optimizing large-scale Spark workloads for performance, scalability, and infrastructure cost efficiencyExperience with containerization and infrastructure technologies such as Docker, Kubernetes, and TerraformExperience working with relational or distributed databases such as PostgreSQL or RedshiftYou bring strong hands‑on engineering expertise across distributed computing, large‑scale data processing, and infrastructure optimization while also helping guide technical direction and mentor engineers across the organizationStrong grasp of software engineering fundamentals including distributed systems, data structures, concurrency, and system designExperience working with healthcare, life sciences, Real World Evidence (RWE), or large‑scale healthcare datasets is strongly preferredStrong coding experience in Python (PySpark), Scala, Java, or equivalent languages used for distributed processing systemsDeep expertise with distributed data processing frameworks such as Apache Spark and Hadoop, particularly within AWS environmentsStrong communication and collaboration skills across both technical and non‑technical stakeholdersExperience improving performance, scalability, observability, and infrastructure efficiency within distributed systemsExperience with streaming and event‑driven architectures using technologies such as Kafka or Spark StreamingExperience with entity resolution, identity mapping, automatching, deduplication, or large‑scale matching systems is strongly preferredProven ability to operate effectively within highly scalable, production‑grade distributed systemsFamiliarity with modern development and infrastructure tooling including Git, CI/CD pipelines, Docker, Kubernetes, Terraform, Argo, Hudi, and JIRA8+ years of experience building and maintaining large‑scale distributed data systems and pipelinesExperience with streaming technologies such as Kafka, Spark Streaming, or KSQLExperience performing root cause analysis across large‑scale distributed systems and complex data pipelinesStrong proficiency in Python (PySpark), Scala, Java, or other modern programming languages used for large‑scale distributed processingDemonstrated technical leadership experience mentoring engineers and driving complex technical initiativesExperience with orchestration and lakehouse technologies such as Argo and Hudi or comparable platformsYou are an experienced data engineer with deep expertise building and optimizing distributed data systems in cloud‑native environments. You thrive solving complex scalability and performance challenges across high‑volume data processing systems and enjoy operating in highly technical, fast‑paced engineering environmentsAbility to write clean, maintainable, modular, and production‑grade codeStrong understanding of distributed file formats including Apache Parquet and Apache AVRO