Principal Data Scientist (Founding)

Katalyze AI, Inc.

Toronto

On-site

CAD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Katalyze AI, Inc. is seeking a Principal Data Scientist to work at the nexus of applied statistics, machine learning, and biotechnology. You will analyze complex data, build interpretable models, and provide actionable recommendations for biopharma and manufacturing customers.

This high-ownership, customer-facing role collaborates with scientists and engineers at client accounts to deliver end-to-end data science solutions and robust decision support across platforms.

Qualifications

  • 6+ years of applied data science experience with statistics and ML
  • Strong understanding of model families: linear/logistic regression, tree-based models, SVMs, neural networks, transformers
  • Strong time series expertise: Fourier analysis, wavelet transforms, forecasting
  • Rigorous model evaluation: train/test design for time-series, SHAP, uncertainty quantification
  • Experience with Gaussian processes, Bayesian methods, or uncertainty quantification
  • Strong Python skills: scikit-learn, pandas, numpy, statsmodels, PyTorch or TensorFlow
  • Experience with enterprise data infrastructure — SQL, data warehouses, cloud platforms (Snowflake, Databricks, Redshift)
  • Excellent communication to scientists, engineers, and business stakeholders
  • Experience with scientific/industrial data is a plus
  • LLMs or agentic systems a plus, not core requirement

Responsibilities

  • Build predictive and diagnostic models on scientific and industrial data
  • Choose applicable modelling techniques with justified reasoning
  • Apply signal processing and time series methods to sensor and process data
  • Design robust model evaluation frameworks for time-series data
  • Create interpretable ML pipelines that surface drivers of variability
  • Develop analytics dashboards to communicate findings to technical and non-technical stakeholders
  • Collaborate with deployment teams to deliver data science components for customers
  • Engage with enterprise customers to understand data challenges and quality systems
  • Explore LLM-based approaches to automate insights and analytics workflows

Skills

Python
Time series analysis
Statistics
ML modeling
Data visualization

Education

PhD or Master’s in Data Science/Statistics
Chemical Engineering or related field

Tools

XGBoost
LightGBM
CatBoost
SQL
Snowflake
Databricks
Redshift

Job description

About Katalyze AI

Katalyze AI is a fast-growing AI-driven biotech platform company on a mission to make life-saving drugs accessible and affordable for everyone. Our AI Agents help pharmaceutical and biotech companies increase production efficiency, reduce costs, and minimize waste. We're a team of humble, fast-moving, and curious craftspeople working at the intersection of science and AI.

About the Role

We're looking for a Principal Data Scientist to join Katalyze AI and work at the intersection of applied statistics, machine learning, and biotechnology. You'll independently analyze complex scientific and process data build interpretable predictive models, and translate findings into actionable recommendations for enterprise customers in biopharma and advanced manufacturing.

This is a high-ownership, customer-facing role. You'll work directly with scientists and engineers at our accounts, not just hand off reports internally.

What You'll Do

Build predictive and diagnostic models on scientific and industrial data (time series, multivariate sensor data, spectral data, batch records)

Select and apply the right modelling technique for each problem — gradient-boosted trees, Gaussian processes, neural networks, classical statistical models — with clear reasoning for your choices

Apply signal processing and time series methods (Fourier transforms, wavelet analysis, autocorrelation, decomposition, forecasting) to real-world sensor and process data

Design rigorous model evaluation frameworks: cross-validation strategies for time-series data, SHAP-based interpretability, uncertainty quantification, and statistical significance testing

Build interpretable ML pipelines that surface drivers of variability in ways that satisfy audit and documentation requirements

Design analytics dashboards that communicate complex statistical findings to manufacturing scientists, quality teams, and supply chain managers

Work closely with the Deployment Strategist to configure and deliver data science components for customer deployments

Partner directly with enterprise customers to understand their data challenges, deviation patterns, and quality systems

Apply LLM-based approaches where appropriate to automate insight generation and multi-step analytical workflows

What We're Looking For

6+ years of applied data science experience with a strong foundation in statistics and machine learning

Deep understanding of how model families work — linear/logistic regression, tree-based models (XGBoost, LightGBM, CatBoost), SVMs, neural networks, transformers — and when to use each

Strong time series expertise: Fourier analysis, wavelet transforms, autocorrelation, stationarity, decomposition, and forecasting (ARIMA, Prophet, and deep learning approaches)

Rigorous model evaluation skills: proper train/test design for time-series data, overfitting detection, SHAP and interpretability methods, uncertainty quantification

Experience with Gaussian processes, Bayesian methods, or uncertainty quantification

Strong Python skills: scikit-learn, pandas, numpy, statsmodels, PyTorch or TensorFlow

Experience with enterprise data infrastructure — SQL, data warehouses, cloud platforms (Snowflake, Databricks, Redshift)

Strong communication skills — able to explain statistical findings clearly to scientists, engineers, and business stakeholders

Experience with scientific or industrial data (sensor streams, spectral data, batch records, LIMS/MES outputs) is a strong plus

PhD or Master's in Data Science, Statistics, Chemical Engineering, or related field preferred

Experience with LLMs or agentic systems is a plus, not a core requirement

ML & Analytics: Python, scikit-learn, XGBoost, LightGBM, PyTorch, statsmodels

LLM / Agents: Claude/GPT APIs, LangChain (where applicable)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Data Engineer
Staff Data Engineer

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Senior ML Scientist
Senior ML Scientist

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Forward Deployed Engineer (Staff/ Founding)
Forward Deployed Engineer (Staff/ Founding)

Katalyze AI, Inc. • Toronto

On-site
CAD 140,000 - 180,000
Data Science Expert
Data Science Expert

CoFoMo Inc. • Montreal (administrative region)

On-site
CAD 120,000 - 180,000
Growth Marketing (Founding)
Growth Marketing (Founding)

Katalyze AI, Inc. • Toronto

On-site
CAD 120,000 - 180,000
Account Manager (Mid-Market)
Account Manager (Mid-Market)

Katalyze AI, Inc. • Toronto

Hybrid
CAD 65,000 - 90,000
Senior Product Designer
Senior Product Designer

Katalyze AI • Toronto

On-site
CAD 80,000 - 100,000
Data Scientist
Data Scientist

Charger Logistics Inc • Brampton

On-site
CAD 90,000 - 150,000
Competitive Salary
Healthcare Benefit Package
Career Growth
Data Analyst
Data Analyst

Astreya • Toronto

On-site
CAD 90,000 - 130,000
Founding Engineer (Agentic Platform)
Founding Engineer (Agentic Platform)

Katalyze AI, Inc. • Toronto

On-site
CAD 130,000 - 200,000