Staff Data Scientist

Vi

Boston (MA)

On-site

USD 140,000 - 220,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Vi Engage is seeking a senior data scientist to own scalable ML pipelines training, scoring, and deployment across health system deployments. You will turn longitudinal data into predictive models that generalize across customers with automated feature engineering and model selection.

Role emphasizes production quality, end-to-end pipeline building in Python stack (pandas, sklearn, PySpark, Airflow) and cloud tooling (AWS/SageMaker). On-site Boston preferred.

Qualifications

  • Shipped uplift or propensity models in production that changed business outcomes.
  • POC to product: repeatable deployments across multiple customers.
  • Deep data science: segmentation, campaign optimization, uplift estimation.
  • Proficient in Python stack and end-to-end pipeline construction.

Responsibilities

  • Develop and maintain scalable ML pipelines that train, score, and deliver predictions.
  • Implement automatic feature engineering and model selection for customer deployments.
  • Improve cross-customer generalization across workloads.
  • Productize pilots into durable capabilities for future deployments.

Skills

Python
ML engineering
Pandas
PySpark
Airflow
SageMaker

Tools

SageMaker
PySpark
Airflow
Pandas
SQL

Job description

Vi Engage puts predictive models into the daily operations of the largest health systems and health plans in the country — driving care navigation, specialty capture, and the workflows that follow from them. This role builds the modeling engine those deployments run on.

You will own the pipelines that turn longitudinal claims, EHR, lab, and online behavioral intent data into predictions about which patients face upcoming healthcare utilization and which are high-propensity to enroll in preventative care programs. Not one model for one customer — the config-driven machinery that trains, selects, scores, and delivers models for any customer, with automatic feature engineering and model selection doing the work that bespoke engineering does today.

The measure of this work is generalization. A model that lifts enrollment for one health plan is a good result; a pipeline that reproduces that lift for the next twenty without per-customer engineering is the product. You set the modeling standard that Forward Deployed Data Scientists build their customer deployments against, and you are accountable for the improvements that hold across all of them.

What You'll Own
  • ML pipelines; Built for scale. Config-driven pipelines that train, score, and deliver predictions from large longitudinal datasets.
  • Mass customization. Automatic feature engineering and model selection that produces a customer-specific model without customer-specific work.
  • Cross-customer generalization. Find the modeling improvements that hold up across every deployment.
  • Building blocks for the field. Implementing and maintaining the DS components that Forward Deployed Data Scientists compose into running customer deployments, so the next deployment is faster than the last.
  • Productization. Turn pilots and one-off proofs into product capabilities that survive their fourth customer.
What We're Looking For
  • Modeling longitudinal data in production. You have personally shipped uplift, survival, or propensity models whose output changed how an organization spent money or who it reached. This is the capability we screen hardest on.
  • POC to product. You have taken a pilot or proof of concept and turned it into something repeatable that survived a second, third, and fourth customer. You know which parts of a one-off are the product and which parts are the customer.
  • Real data science depth. Segmentation, campaign optimization, and the identification strategy behind an uplift estimate. You can defend a modeling choice to a skeptical internal analytics team, evaluate a model honestly, and say when a simpler approach is the right answer.
  • Python and ML engineering. Fluent in Python and the working stack — pandas, sklearn, PySpark, airflow. You build the pipeline, not just the model inside it.
  • MLOps. Model tracking and deployment tooling — mlflow, SageMaker, or comparable
  • AWS and Cloud. You don’t need to hand-off to a dedicated engineer. You can get your models running at scale in the cloud using our AWS stack: S3, Glue, EMR, MWAA, SageMaker.
Nice To Have
  • Healthcare or life sciences domain knowledge — claims, EHR, HL7/FHIR, lab data, or population health analytics
  • Familiarity with HIPAA and healthcare compliance and data governance frameworks
  • Experience designing pilots and efficacy studies that tie model performance to a business outcome
  • Experience building internal platforms or frameworks that other engineers build on top of
What This Role Is Not

This is an applied, in-production role. It is not a research position — the work is measured by pipelines that run and models that hold up across customers, not by novelty. It is primarily not a customer-facing role (Applied Data Scientists own the customer accounts) but you may interface with design partner clients on occasion. This is not a management role — you will be hands-on architecting and building this product with the team.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Data Scientist
Staff Data Scientist

Getvi • Boston (MA)

On-site
USD 180,000 - 240,000
Applied Data Scientist
Applied Data Scientist

Getvi • Boston (MA)

On-site
USD 180,000 - 230,000
Applied Data Scientist
Applied Data Scientist

Vi • Boston (MA)

On-site
USD 120,000 - 180,000
Data Scientist
Data Scientist

Stealth Startup • New York (NY)

On-site
USD 200,000 - 225,000
Relocation package
Competitive equity
Marquee benefits package
Senior Data Scientist
Senior Data Scientist

Synapse Health • United States

On-site
USD 120,000 - 190,000
Professional growth opportunities
Flexible PTO
Medical, dental, vision, STD & LTD
+1
Forward Deployed Data Engineer
Forward Deployed Data Engineer

Total Performance Consulting • United States

Hybrid
USD 150,000 - 190,000
Staff Data Scientist - Scalable ML Pipelines for Healthcare
Staff Data Scientist - Scalable ML Pipelines for Healthcare

Vi • Boston (MA)

On-site
USD 140,000 - 220,000
Data Scientist Lead
Data Scientist Lead

Kontakt Micro-Location Sp. Z.o.o. • New York (NY)

Hybrid
USD 180,000 - 240,000
Health benefits
401(k)
Paid time off
+2
Data Science Tech Lead
Data Science Tech Lead

Univedge Consulting LLC • Minneapolis (MN)

On-site
USD 180,000 - 240,000
Senior Data Scientist
Senior Data Scientist

DrFirst • Northern (KY)

Hybrid
USD 140,000 - 170,000
Discretionary bonus
Medical, dental, and vision insurance
401K with match
+3