The Data Scientist turns CCAD's clinical, operational, imaging research, digital-front-door, and device data into evidence-based insights and validated AI solutions that advance patient care quality, safety, efficiency, and discovery. The role leads problem framing, cohort and feature design, statistical and causal analysis, experimentation, predictive modelling, computer vision, and multimodal AI evaluation, and clear communication of findings to clinical and business stakeholders. As CCAD advances toward the north star of an autonomous hospital, the Data Scientist partners with clinicians, clinical informatics, data engineering, AI/ML engineering, governance, and product teams to convert high-value use cases into safe, measurable, human-in-the-loop analytics and AI products.
What You’ll Do
Primary Job Function
- Partner with clinical, operational, and digital teams to prioritize AI use cases aligned to autonomous hospital goals such as patient flow, capacity, staffing, demand forecasting, diagnostics, clinical decision support, and command‑center automation.
- Translate clinical and operational problems into measurable hypotheses, data requirements, benefit cases, user stories, success metrics, and safe human-in-the-loop workflows.
- Create analytic, validation, and adoption plans with clinicians, product owners, and governance stakeholders.
Healthcare Data Understanding & Feature Engineering
- Extract, profile, cleanse, and validate structured, semi-structured, and unstructured data from EHR, ERP, scheduling, labs, pharmacy, PACS/RIS/DICOM, IoMT devices, patient experience, contact‑center, and digital channels.
- Build clinically meaningful cohorts and features using Epic Clarity, Caboodle, Cogito, FHIR, HL7, DICOM, DICOMweb, OMOP, SNOMED CT, LOINC, RxNorm, and other terminology standards.
- Apply data‑quality rules, de‑identification, missing‑data handling, imputation, leakage checks, normalization, NLP extraction, and reproducible feature documentation.
Statistical, Causal, Experimental Analysis
- Conduct EDA, hypothesis testing, regression, Bayesian analysis, survival analysis, and time‑series forecasting.
- Perform experimental design, A/B testing, and quasi‑experimental impact evaluation.
- Analyze subgroup performance, equity, fairness, calibration, clinical utility, operational ROI, and uncertainty.
- Produce clear narratives, visuals, and recommendations for technical and non-technical audiences.
Model Development & Validation
- Develop and evaluate supervised, unsupervised, semi-supervised, NLP, forecasting, optimization, simulation, and reinforcement-learning models using modern open-source and cloud ML frameworks.
- Establish baselines, train-validation-test splits, temporal validation, cross-validation, hyper-parameter tuning, error analysis, explainability, and confidence intervals.
- Prepare champion-model documentation including model cards, validation packs, operating thresholds, limitations, expected failure modes, and production handoff materials.
Computer Vision & Multimodal AI
- Build, fine-tune, or evaluate computer-vision models for medical imaging, digital pathology, video-enabled operations, or document imaging.
- Apply CNNs, U-Net, Vision Transformers, object detection, segmentation, classification, MONAI, OpenCV, DICOM/DICOMweb, and multimodal LLMs as appropriate.
- Evaluate clinical and operational usefulness using metrics such as sensitivity, specificity, AUROC, AUPRC, Dice, IoU, calibration, turnaround time, false-alert burden, and human-review requirements.
Generative AI, LLMs & Agentic Workflows
- Prototype and evaluate Retrieval‑Augmented Generation (RAG), summarization, information extraction, classification, conversational analytics, clinical documentation support, and agentic workflow use cases.
- Design prompt, retrieval, and tool-use evaluation sets; test hallucination, grounding, bias, privacy, prompt-injection, and safety risks.
- Work with AI/ML Engineers to define structured outputs, guardrails, approved knowledge sources, monitoring measures, and deployment criteria.
Visualization, Decision Support & Storytelling
- Build dashboards, analytical applications, what-if simulators, and command‑center decision aids using Power BI, Tableau, Streamlit, Shiny, Dash, or Power Apps.
- Design visualizations and explainability views that make model output interpretable and actionable for clinicians, operational leaders, and frontline teams.
- Measure adoption, user feedback, workflow fit, and benefits realization after deployment.
Responsible Clinical AI Governance
- Apply privacy, data governance, clinical safety, responsible AI, human oversight, change control, auditability, and cybersecurity requirements throughout the model lifecycle.
- Create and maintain model cards, data sheets, validation summaries, monitoring thresholds, fairness assessments, and retraining recommendations.
- Participate in AI governance, safety, privacy, and clinical validation reviews before production use.
Collaboration & Agile Delivery
- Work in Agile squads with clinicians, product owners, data engineers, AI/ML engineers, BI developers, and governance teams to deliver measurable outcomes.
- Write reusable, modular, version-controlled analysis and modeling code with peer review and reproducibility practices.
- Support production analysis, incident investigation, and after-hours validation when changes affect critical analytics or AI services.
What You’ll Bring
- Expertise in Python, SQL, self‑hands‑on use of Pandas, NumPy, Polars, PySpark, Spark SQL, scikit-learn, statsmodels, and notebook‑to‑production development practices.
- Strong statistical foundation in probability, Bayesian inference, regression, causal inference, survival analysis, time‑series forecasting, experiment design, sample‑size estimation, power analysis, and uncertainty quantification.
- Advanced ML expertise across supervised, unsupervised, NLP, forecasting, optimization, simulation, and reinforcement-learning methods with appropriate evaluation metrics and clinical interpretation.
- Computer vision expertise for healthcare: DICOM, PACS/RIS patterns, CNNs, U-Net, Vision Transformers, segmentation, object detection, digital pathology, medical imaging, video analytics, and multimodal vision‑language models.
- Generative AI expertise with LLM evaluation, prompt design, RAG, embeddings, vector search, knowledge graphs, structured outputs, agentic workflow design, hallucination grounding, measurement, and safe human-in-the-loop design.
Bachelor’s degree in Statistics, Data Science, Computer Science, Mathematics, Physics, Biomedical Engineering, Epidemiology, Health Informatics, Health Economics or related quantitative field. PREFERRED: Master’s degree or PhD in relevant discipline. 3+ years in data science, advanced analytics, applied ML/AI or applied research, including healthcare, life sciences, hospital operations or similarly regulated data environments. 5+ years in healthcare AI/analytics; experience leading AI products from problem definition through validation and adoption; peer‑reviewed research or clinical quality improvement experience.