Staff Data Scientist

Getvi

Boston (MA)

On-site

USD 180,000 - 240,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Getvi is seeking an applied data scientist/ML engineer to own end-to-end predictive modeling pipelines using longitudinal data from health systems. You will deploy models that identify patients at risk and enable proactive care programs.

In this role, you will scale config-driven pipelines, ensure cross-customer generalization, and work with the cloud stack (AWS) to ship reusable components. Healthcare domain knowledge is a plus.

Qualifications

  • Proven track record shipping uplift, survival, or propensity models that changed spending or reach.
  • Turn pilots or proofs of concept into repeatable production deployments across customers.
  • Fluent Python with pandas, sklearn, PySpark and Airflow.
  • Experience with MLOps, model tracking and deployment tooling.

Responsibilities

  • Own ML pipelines from data to deployed predictions at scale.
  • Build config-driven, reusable components for multiple customers in AWS.
  • Defend modeling choices to analytics stakeholders and justify simpler approaches when needed.
  • Collaborate with engineers to create internal platforms used by other teams.

Skills

Python
ML Engineering
Pandas
Scikit-learn
PySpark
Airflow
MLflow
SageMaker

Tools

AWS S3
AWS Glue
AWS EMR
AWS MWAA
SageMaker

Job description

Role Summary

Vi Engage puts predictive models into the daily operations of the largest health systems and health plans in the country — driving care navigation, specialty capture, and the workflows that follow from them. This role builds the modeling engine those deployments run on.You will own the pipelines that turn longitudinal claims, EHR, lab, and online behavioral intent data into predictions about which patients face upcoming healthcare utilization and which are high-propensity to enroll in preventative care programs. Not one model for one customer — the config-driven machinery that trains, selects, scores, and delivers models for any customer, with automatic feature engineering and model selection doing the work that bespoke engineering does today.The measure of this work is generalization. A model that lifts enrollment for one health plan is a good result; a pipeline that reproduces that lift for the next twenty without per-customer engineering is the product. You set the modeling standard that Forward Deployed Data Scientists build their customer deployments against, and you are accountable for the improvements that hold across all of them.

What You'll Own
  • ML pipelines;
  • Built for scale.
  • Config-driven pipelines that train, score, and deliver predictions from large longitudinal datasets.
  • Mass customization.
  • Automatic feature engineering and model selection that produces a customer-specific model without customer-specific work.
  • Cross-customer generalization.
  • Find the modeling improvements that hold up across every deployment.
  • Building blocks for the field.
  • Implementing and maintaining the DS components that Forward Deployed Data Scientists compose into running customer deployments, so the next deployment is faster than the last.
  • Productization.
  • Turn pilots and one-off proofs into product capabilities that survive their fourth customer.
What We're Looking For
  • Modeling longitudinal data in production.
  • You have personally shipped uplift, survival, or propensity models whose output changed how an organization spent money or who it reached.
  • This is the capability we screen hardest on.
  • POC to product.
  • You have taken a pilot or proof of concept and turned it into something repeatable that survived a second, third, and fourth customer.
  • You know which parts of a one-off are the product and which parts are the customer.
  • Real data science depth.
  • Segmentation, campaign optimization, and the identification strategy behind an uplift estimate.
  • You can defend a modeling choice to a skeptical internal analytics team, evaluate a model honestly, and say when a simpler approach is the right answer.
  • Python and ML engineering.
  • Fluent in Python and the working stack — pandas, sklearn, PySpark, airflow.
  • You build the pipeline, not just the model inside it.
  • MLOps.
  • Model tracking and deployment tooling — mlflow, SageMaker, or comparableAWS and Cloud.
  • You don't need to hand-off to a dedicated engineer.
  • You can get your models running at scale in the cloud using our AWS stack: S3, Glue, EMR, MWAA, SageMaker.
  • Nice To Have
  • Healthcare or life sciences domain knowledge — claims, EHR, HL7/FHIR, lab data, or population health analytics
  • Familiarity with HIPAA and healthcare compliance and data governance frameworks
  • Experience designing pilots and efficacy studies that tie model performance to a business outcome
  • Experience building internal platforms or frameworks that other engineers build on top of
What This Role Is Not

This is an applied, in-production role. It is not a research position — the work is measured by pipelines that run and models that hold up across customers, not by novelty. It is primarily not a customer-facing role (Applied Data Scientists own the customer accounts) but you may interface with design partner clients on occasion. This is not a management role — you will be hands‑on architecting and building this product with the team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Data Scientist
Staff Data Scientist

Vi • Boston (MA)

On-site
USD 140,000 - 220,000
Applied Data Scientist
Applied Data Scientist

Getvi • Boston (MA)

On-site
USD 180,000 - 230,000
Applied Data Scientist
Applied Data Scientist

Vi • Boston (MA)

On-site
USD 120,000 - 180,000
Senior Data Scientist
Senior Data Scientist

Synapse Health • United States

On-site
USD 120,000 - 190,000
Professional growth opportunities
Flexible PTO
Medical, dental, vision, STD & LTD
+1
Forward Deployed Data Engineer
Forward Deployed Data Engineer

Perform • Nashville (TN)

Hybrid
USD 150,000 - 190,000
Forward Deployed Data Engineer
Forward Deployed Data Engineer

Perform • Atlanta (GA)

Hybrid
USD 140,000 - 190,000
Forward Deployed Data Engineer
Forward Deployed Data Engineer

Perform • Florida City (FL)

Hybrid
USD 130,000 - 185,000
Forward Deployed Data Engineer
Forward Deployed Data Engineer

Perform • Los Angeles (CA)

Hybrid
USD 170,000 - 210,000
Hybrid work model
Health benefits
Data Scientist
Data Scientist

Stealth Startup • New York (NY)

On-site
USD 200,000 - 225,000
Relocation package
Competitive equity
Marquee benefits package
Staff Data Scientist
Staff Data Scientist

Synapse Health • United States

On-site
USD 180,000 - 240,000
Health insurance
401(k) with company match