Senior Data Engineer

ParetoHealth

Deutschland

Hybrid

EUR 121.000 - 164.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Fully paid medical, dental, and vision
Flexible PTO
401k company contribution
Tuition reimbursement
Professional development allowance
Transportation allowance and parking
Engaging hybrid work environment

Zusammenfassung

ParetoHealth is hiring a Senior Data Engineer to design, build, and maintain serverless data pipelines on AWS, enabling analytics and AI workflows.

You will collaborate with Data Science to support MLOps, feature stores, and robust data contracts. The role emphasizes scalable data infrastructure, data governance, and cost-efficient operations in a hybrid work setting.

Qualifikationen

  • Minimum 5+ years in data engineering or related role.
  • Experience with AWS data services in production environments.
  • Strong proficiency in Python and SQL; PySpark preferred.

Aufgaben

  • Design, build, and maintain serverless data pipelines on AWS.
  • Develop and evolve data models and feature stores for AI/Analytics teams.
  • Collaborate with Product, Analytics, and AI stakeholders to translate requirements into robust data structures.
  • Implement data quality controls, monitoring, and governance for sensitive data.
  • Optimize serverless workloads for cost, performance, and scalability.
  • Contribute to CI/CD and Infrastructure as Code practices (AWS CDK).
  • Support ML model training and scoring pipelines in collaboration with Data Science.

Kenntnisse

AWS
Python
SQL
PySpark
TypeScript
MLOps
Data modeling
CI/CD
Kedro
Data governance

Ausbildung

Bachelor's degree in CS/Engineering/Data Science
Master's degree preferred

Tools

AWS Lambda
AWS Glue
AWS Athena
S3
Kafka
Kinesis
AWS CDK

Jobbeschreibung

About ParetoHealth

ParetoHealth is redefining the way employers fund healthcare. As the largest and fastest-growing benefits captive in the United States, we help thousands of small and midsize employers take control of healthcare costs through a smarter, more sustainable model.

Our mission is simple: give small and midsize employers the scale and protection they need to eliminate volatility and lower healthcare costs.

By combining data-driven insights, innovative risk management, and the collective purchasing power of our community, we enable employers to reduce volatility, improve long-term outcomes, and reinvest savings into their businesses and their people.

Headquartered in Philadelphia, ParetoHealth is growing rapidly and transforming one of the country's largest industries. Our success is fueled by talented people who are united by our four core values: Fire in the Belly, For the Greater Good, See the Field, and Get It Done Right. These values shape how we innovate, collaborate, make decisions, and deliver exceptional results for our clients and one another.

If you're energized by solving complex challenges, thrive in a high-growth environment, and want to help reshape the future of healthcare, we'd love to meet you.

Please note that ParetoHealth does not provide employment visa sponsorship for this position. Candidates must be authorized to work in the United States without sponsorship both now or in the future.

Position Summary:

The Senior Data Engineer will design, build, and maintain serverless data pipelines and data models on AWS that make high-quality, analytics-ready data available to our AI and Analytics teams. This role will own ingestions, transformation, and storage patterns using services such as Lambda, Glue, Athena, and S3, ensuring data is reliable, well-documented, and aligned with business and model-training needs. The role will also partner with Data Science to support MLOps, including reproducible machine-learning training and scoring pipeline, automated feature generation, deployment, and monitoring. It will power ML model workflows using sensitive healthcare data and support model training and scoring, production monitoring, and business feedback loops.

Key Responsibilities:
  • Design, implement, and maintain scalable, fully serverless data pipelines on AWS using Lambda, Glue, Athena, Step Functions, and S3 to support reporting, analytics, and AI use cases.
  • Build and evolve data models and schemas that enable performant querying and downstream consumption by AI, Analytics and engineering teams. Build versioned data models, feature stores, with semantics, reproducible backfills, and safeguards against leakage or inconsistent definitions.
  • Develop ETL/ELT workflows to ingest, cleanse, transform, and load data from internal applications, third‑party sources, and event streams into our data lake and analytical layers.
  • Partner closely with Product, Underwriting, Analytics, and AI stakeholders to understand data requirements and translate them into robust data structures, contracts, and SLAs.
  • Implement data quality controls, monitoring, and alerting to ensure accuracy, completeness, timeliness, and lineage of critical datasets and features used by models. Include automated controls for schema change, validity, reconciliation, claims maturity, and model leakage.
  • Optimize serverless workloads for cost, performance, and scalability, including query tuning in Athena and efficient storage formats/partitioning in S3.
  • Contribute to and enforce data engineering best practices, including version control, code review, CI/CD for data pipelines, and Infrastructure as Code with AWS CDK.
  • Collaborate with AI and Analytics teams to design and maintain feature stores and other reusable data assets that accelerate experimentation and model deployment. Partner on batch or API scoring and capture model/data versions, recommendations, actions, overrides, and claims to close the learning loop.
  • Experience designing and operating event-driven or streaming data pipelines for near-real-time scoring or decisioning, using Kafka, Amazon Kinesis, or comparable technologies; strong understanding of event schemas, idempotency, ordering, retries, observability, and failure recovery.
  • Troubleshoot pipeline issues, resolve data‑related incidents, and provide ongoing support for production data workflows and model‑driven applications.
  • Document data models, pipelines, and data contracts, and help evangelize data literacy and self‑service analytics across the organization.
Required Skills & Qualifications:
  • Strong experience with AWS data and serverless services (Lambda, Glue, Athena, S3, Step Functions or similar orchestration tools).
  • Advanced proficiency in Python (required) and SQL, with experience in relational and NoSQL databases and in building reusable ETL/ELT libraries. Hands-on experience with PySpark for distributed data processing is strongly preferred; TypeScript experience for AWS CDK required.
  • Experience with MLOps practices, including versioned feature and training datasets, orchestration of training and scoring workflows, CI/CD, model/data monitoring, and reproducible deployments in partnership with AI.
  • Solid understanding of data modeling principles for AI workloads. Apply schema evolution, incremental processing, partitioning, and columnar or open-table formats such as Parquet and Iceberg.
  • Experience designing and operating data pipelines at scale, including batch and near‑real‑time ingestion. Hands‑on experience with Kafka is preferred.
  • Familiarity with SQL and query optimization in columnar data stores and engines.
  • Knowledge of data quality, governance, and security for sensitive healthcare and financial data, including least‑privilege access, retention, classification, tokenization, and de‑identification.
  • Hands‑on experience with Infrastructure as Code, preferably AWS CDK, for provisioning and managing data infrastructure.
  • Ability to partner with AI and Analytics teams to design data solutions for model experimentation and production; Kedro or similar workflow experience is a plus.
  • Knowledge of responsible AI and data‑use practices, including permitted‑use controls, privacy, fairness considerations, documentation, and governance for sensitive healthcare and financial data.
  • Strong communication skills to explain complex data concepts clearly to business stakeholders; regulated‑domain experience, preferably healthcare claims or insurance, is a plus.
Minimum Requirements:
  • Bachelor's degree in Computer Science, Engineering, Mathematics, Data Science, or a related field, or equivalent practical experience; a master's degree is preferred.
  • 5+ years in data engineering or a related role, with significant ownership of production data pipelines and analytics platforms; hands‑on ML/AI experience supporting feature preparation, model training, deployment, or monitoring is strongly preferred.
  • Experience with AWS data services in a production environment (for example, Lambda, Glue, Athena, and S3) is preferred; equivalent production experience on another cloud data platform will be considered.
  • Experience building and maintaining data solutions that support reporting, BI, and/or AI/ML initiatives.

Perks & Benefits:

  • Fully paid medical, dental, and vision benefits.
  • Flexible PTO
  • 401k company contribution
  • Tuition reimbursement
  • Professional development allowance
  • Transportation allowance and daily parking reimbursement
  • Engaging hybrid work environment
We are guided by our values:

Fire in the belly

The drive to learn, to improve, and to deliver outstanding value everyday.

See the field

The ability to see the big picture and prepare to meet tomorrow’sneeds.

Get it done right

The passion to produce at higher rates and to the higheststandards.

For the greater good

A united community creating better health benefit solutions forall.

Please note that any communication from our recruiters and hiring managers at ParetoHealth about a job opportunity will only be made by a ParetoHealth employee with an @paretohealth.com address. ParetoHealth does not conduct text message or chat-based interviews. Any other email addresses, agencies, or forums may be phishing scams designed to obtain your personal information.
We will not ask you to provide personal or financial information, including, but not limited to, your social security number, online account passwords, credit card numbers, passport information, and other related banking information until we begin onboarding activities, which will be coordinated by a member of the ParetoHealth People Ops Team with an @paretohealth.com email address.
Disclosures:

ParetoHealth is an Equal Opportunity Employer and does not discriminate on the basis of race, color, religion (creed), gender, gender expression, age, national origin (ancestry), disability, marital status, sexual orientation, or military status, in any of its activities or operations. These activities include, but are not limited to, hiring and firing of staff, selection of volunteers and vendors, and provision of services. We are committed to providing an inclusive and welcoming environment for all members of our staff, clients, volunteers, subcontractors, vendors, and clients.

California Applicants: See Pareto’s CCA Notice of Collection for California Employees and Applicants for information about how Pareto Captive Services, LLC, Pareto Health, LLC, and Pareto Underwriting Partners, LLC, together with their respective subsidiaries (collectively, “Pareto”) collects and uses personal information submitted by employment applicants.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Software Engineer
Senior Software Engineer

ParetoHealth • Deutschland

Hybrid
EUR 121.000 - 147.000
Fully paid medical, dental, and vision
Flexible PTO
401k company contribution
+3
Senior Software Engineer
Senior Software Engineer

paretocaptiveservicesllc • Deutschland

Hybrid
EUR 104.000 - 147.000
Fully paid medical, dental, and vision
Flexible PTO
401k company contribution
+4
QA Automation Engineer
QA Automation Engineer

ParetoHealth • Deutschland

Hybrid
EUR 78.000 - 104.000
Fully paid medical, dental, and vision
Flexible PTO
401k company contribution
+4
Staff Data Scientist
Staff Data Scientist

Pivotal Health • Deutschland

Hybrid
EUR 110.000 - 160.000
Competitive compensation including equ
Health, dental, and vision coverage
401(k) retirement plan
+2
Director Data Management & Intelligence
Director Data Management & Intelligence

GoHealth Urgent Care • Deutschland

Hybrid
EUR 120.000 - 190.000
Senior Data Engineer (Internal Platform)
Senior Data Engineer (Internal Platform)

Anyline • Berlin

Vor Ort
EUR 70.000 - 90.000
Staff / Principal Forward-Deployed Architect, Data Modernization
Staff / Principal Forward-Deployed Architect, Data Modernization

Transformcap • Deutschland

Hybrid
EUR 164.000 - 217.000
Equity packages
Robust medical/dental/vision insurance
Flexible working hours
+1
Senior Data Engineer
Senior Data Engineer

jobr.pro • Deutschland

Hybrid
EUR 60.000 - 90.000
Generous equity grants
Flexible PTO
401(k) match
(Canada) Sr. Data Governance Analyst, 6 month contract
(Canada) Sr. Data Governance Analyst, 6 month contract

Embedded Shishya • Deutschland

Hybrid
EUR 78.000 - 114.000
Benefits starting from Day 1
Senior Data Engineer (Backend Development Experience) (f/m/d)
Senior Data Engineer (Backend Development Experience) (f/m/d)

myoncare • München

Vor Ort
EUR 80.000 - 100.000
Competitive salary with performance-based growth
Central located office with lunch options
Global team collaboration
+3