If you’ve built real, production-grade AWS Glue pipelines — not just experimented with them — this is the kind of role that lets you own it properly.
We’re partnering with a growing Healthcare / PBM (Pharmacy Benefit Management) organization investing heavily in its AWS data platform. Glue is the backbone. This isn’t maintenance work — it’s platform-level engineering handling large-scale pharmacy and claims data in a secure, HIPAA-aware environment.
What You’ll Be Doing:
- Architect and optimize AWS Glue ETL pipelines (Jobs, Crawlers, Workflows, Data Catalog)
- Build and maintain scalable S3 data lake structures
- Design and tune Redshift warehouse environments
- Implement efficient partitioning and performance optimization strategies
- Handle large-scale pharmacy and claims datasets
- Drive data quality, governance, and reliability standards
- Partner with analytics, clinical, and product stakeholders to deliver trusted data assets
- Influence architecture decisions and long-term platform direction
Core Tech Stack:
- PySpark / Spark
- Amazon S3
- Amazon Redshift
- Python & SQL
- Infrastructure-as-Code exposure (CDK or Terraform a strong plus)
What We’re Looking For:
- 5+ years in Data Engineering
- Strong Spark/PySpark capability
- Experience working with large, regulated datasets
- Background in Healthcare, PBM, payer, or insurance data environments preferred
- Comfortable owning architecture — not just executing tickets