Senior Data Scientist

Lemnis

United States

Remote

USD 150,000 - 165,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Mainstay, a division of Lemnis, is expanding its data science team with a Senior Data Scientist role. This position is primarily remote within the United States and carries a salary range of $150,000–$165,000 DOE.

You will develop predictive models, analyze unstructured data, and translate insights into actions for partners and learners. You will own modeling, NLP, and AI tooling, collaborating with data engineers to build scalable, production-ready solutions.

Qualifications

  • 5+ years building predictive models used in production.
  • Practical NLP experience: text classification, embeddings, or similar.
  • Strong SQL as a primary tool for modeling and analysis.
  • Proficient Python for modeling and analysis.
  • Solid applied statistics with appropriate method choice.
  • Experience with modern cloud data warehouses (Snowflake/BigQuery/Databricks).
  • Excellent written and verbal communication; ability to translate findings for non-tech stakeholders.
  • Comfort with ambiguity and prioritization in a fast-paced environment.

Responsibilities

  • Find and ship predictive signals across data sources and partners.
  • Analyze conversational and unstructured data to surface themes and insights.
  • Own AI data tooling and evaluation, including prompts and semantic layers.
  • Build and maintain production-ready models and monitoring workflows.
  • Collaborate with Senior Data Engineer and Analytics Engineer to scale solutions.
  • Document reasoning and tradeoffs in internal knowledge bases.

Skills

Predictive Modeling
NLP
SQL
Python
Statistical Analysis
Communication
Productionization
Production Workflows

Tools

Snowflake
BigQuery
Databricks
Looker
dbt

Job description

Position: Senior Data Scientist
Position Type: Full Time (Primarily Remote)
Salary: $150,000-$165,000 DOE
Join our team at Mainstay, a division of Lemnis! Lemnis is a public charity dedicated to harnessing transformative change to expand learning for all. We are excited to announce a new opportunity for a Senior Data Scientist to join our innovative, forward-thinking, and growing organization. At Lemnis, we are committed to fostering a collaborative and inclusive work environment where every team member can thrive. If you are passionate about expanding learning for all, eager to make a meaningful impact, and ready to take on new challenges, we would love to hear from you.
Position Summary
At Lemnis, we believe that the world is changing in exciting ways. It s up to us to create more equitable, flexible, and learner-centered systems that empower young people to rise to the challenges and opportunities of the future. Do you have questions? At Mainstay, our users have millions, so it s imperative that our systems are stable, robust, and scalable. As a Senior Data Scientist at Mainstay, you will help implement features that help our partners and end users.
About Mainstay
At Mainstay, we believe one conversation can spark a brighter future. Our Engagement Platform makes it easy for colleges and businesses to start and measure conversations that drive action at scale. From our rigorous research methods to our Behavioral Intelligence framework, everything we do is designed to help people take the next step toward achieving their goals.
This is your chance to drive impact at Mainstay and for our users by building and enhancing the systems that power our product. We ll provide the environment for you to master your skills and find both personal and professional growth. We are invested in promoting from within and providing the support and mentorship aimed at your long-term success.
The engineering team is a group of talented full-stack developers, data engineers, analytics engineers, and product experts. Engineers collaborate closely with their counterparts in the product team and external stakeholders to build new functionality and enhance existing products to delight our partners and end users. The team is excited to continue tackling new challenges and leverage cutting edge technologies to solve impactful problems.
About the Role
We r hiring a Senior Data Scientist to find the predictive signals in our data and turn them into insights our partners can act on.
Your early work is less about squeezing out marginal accuracy and more about finding signals. You ll build student-level predictors that sync into partner systems, analyze conversational data at scale to make our AI smarter, improve access to reliable data and insights, and leverage internal tools to scale your expertise.
As the senior data scientist on a lean team, you ll have real say in what we build. We have more promising directions than capacity, so part of the job is deciding which signals are worth pursuing, which analyses will generalize, and which requests to decline. You ll set the methodology bar for modeling and evaluation work here, and you ll be the person others come to when they aren t sure whether to trust a number.
You ll work alongside our Senior Data Engineer and Senior Analytics Engineer. They own infrastructure, pipelines, and the shared semantic layer. You own predictive modeling, unstructured data analysis, AI evaluation, and the analytics surfaces that make data accessible and trustworthy. If you r looking to train models full-time, this isn t that role.
What You ll Do

Find and ship predictive signals

  • Investigate signals across student engagement, outcomes, and partner health, prioritizing by business impact over technical interest.
  • Build predictive models that reach partners through the systems they already use, where output drives real action.
  • Set thresholds against the alert volume teams can actually act on, and be explicit about the cost of a false positive when predictions reach a partner.
  • Evaluate performance across student populations, not just in aggregate, and document known limitations alongside the model.
  • Monitor deployed models for drift, and retire models that stop earning their place.
  • Know when not to build a model. Some questions are better answered with an analysis, a definition change, or a conversation.
  • Own scheduling and monitoring for your models, using orchestration the whole team can maintain.

Analyze our conversational and unstructured data

  • Apply embeddings, clustering, and classification to conversational data, support and service records, and other unstructured sources to surface themes, gaps, and emerging concerns.
  • Turn what you find into changes that improve the product and the partner experience.
  • Build the text analysis foundations that make our unstructured data retrievable and useful to AI tooling.

Own our AI data tooling and evaluation

  • Own our in-warehouse AI configuration, including verified queries, prompts, and agent tooling, and the semantic views that expose your model output. You ll partner with an Analytics Engineer where this depends on the shared semantic layer.
  • Build and maintain the AI evaluation framework for data team work: rubrics precise enough that two reviewers agree, a defensible sampling approach, and reporting that shows whether a model or prompt change actually improved anything.
  • Set the evaluation standard for data team work, and partner with Product and Engineering to share evaluation methodology more broadly.

Make your work reusable

  • Equip internal teams with the data and analysis they need for our most strategic partners, favoring work that generalizes over one- off requests.
  • Grow into building the tooling and training that lets those teams answer questions without the data team.
  • Document reasoning, assumptions, and tradeoffs in our internal data knowledge base as part of finishing the work.
  • Work in dbt/code alongside our engineers, contributing models and requesting changes to shared definitions.
What We r Looking For

Required experience and skills

  • 5+ years building predictive models that someone actually used, including a few years where you owned the problem rather than being handed it, with the judgment that goes with it: calibration, threshold-setting, and knowing when a feature is leaking the answer.
  • Practical NLP experience: text classification, clustering, embeddings, or similar applied work. Calling an LLM API is useful but isn t the same thing.
  • Strong SQL as a primary tool, not a way to get data into a notebook.
  • Working Python for modeling and analysis.
  • Solid applied statistics, with the judgment to know which method fits the question and when the data can t support a conclusion.
  • Modern cloud warehouse experience (Snowflake, BigQuery, Databricks, or similar), especially with in-warehouse AI or agent tooling.
  • Excellent written and verbal communication. You can hand a finding to a non-technical colleague and have them act on it, including knowing what would change your conclusion.
  • Care about how predictions get used. Our scores influence how students get supported, so we want someone who checks whether a model works as well for part-time students as for everyone else, and says so when it d not.
  • Comfort with ambiguity and honesty about uncertainty. We d rather hear “the data can t answer this” than a confident answer that falls apart later.
  • Comfortable working in version control with code review, so your analysis and models are reproducible by someone else.
  • Experience productionizing model output into an operational workflow.
  • Experience scheduling and monitoring recurring jobs in production, and the judgment to reach for tooling the team can maintain rather than a specialized stack that only you know.
  • Track record of choosing what to work on. You ve turned an ambiguous business goal into a scoped project, made the prioritization case, and been accountable for whether it mattered.

Nice to have

  • Transformation tooling (dbt or similar), dimensional modeling, or analytics engineering exposure.
  • Familiarity with AI evals or prompt evaluation.
  • A modern BI tool (Sigma, Looker, Hex, Tableau, or similar) for making findings usable by others.
  • Experience working closely with analytics or data engineers, where your models depended on someone else s tables.
  • Linguistics or computational linguistics background for intent classification and evaluation work.
  • EdTech, higher education, or student success background.

This probably isn t the right opportunity for you if

  • You need a dedicated ML platform to be effective. We deliberately don t run one: models run inside our cloud warehouse, features come from our transformation layer, predictions land in tables that downstream systems read.
  • Your experience is primarily research or model development without anyone using the output.
  • You want to specialize. The role ranges across modeling, text analysis, evaluation, and enablement, and none of them will be someone else s job.
  • You d rather have your work reviewed rather than review others. As the senior person in this discipline here, you ll be setting the standard, not inheriting one.
  • You d introduce a new tool for every problem. We optimize for delivering business value using long-term maintainable solutions, which sometimes means the second-best tool.

Key Performance Metrics

  • Validated predictive signals shipped and in active use by partners or internal teams.
  • Model quality in production terms: calibration and precision at an actionable alert volume.
  • Adoption of AI data tooling, including share of questions answered without data team involvement.
  • Evaluation coverage: share of data team AI surfaces with an active eval, and whether model or prompt changes ship with evidence.
  • Stakeholder feedback from Partner Success, Product, and Leadership.
  • Quality of prioritization: whether the work you chose turned out to matter, and whether you surfaced tradeoffs early.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Founding Data & Machine Learning Lead
Founding Data & Machine Learning Lead

Rippling, Inc. • Bethesda (MD)

Hybrid
USD 126,000 - 154,000
Equity in early-stage company
Health insurance coverage for employee
Dependent coverage (50%) and dental/视觉
Founding Data & Machine Learning Lead
Founding Data & Machine Learning Lead

Rippling, Inc. • United States

Remote
USD 119,000 - 161,000
Equity stake
Health insurance fully covered
Dental & Vision 50% covered
+1
Staff Applied AI Scientist
Staff Applied AI Scientist

Order.co • United States

Remote
USD 180,000 - 260,000
AI Research Scientist, Learning & Evaluation
AI Research Scientist, Learning & Evaluation

Studyfetch • Beverly Hills (CA)

On-site
USD 150,000 - 210,000
Medical, Dental, Vision (100% employer
75% dependent coverage
401(k) with employer matching
+2
Senior Consultant, AI/ML Engineer
Senior Consultant, AI/ML Engineer

Hollstadt Consulting • Minnesota

On-site
USD 150,000 - 210,000
Data Scientist
Data Scientist

ScienceLogic • United States

On-site
USD 120,000 - 180,000
AI Data Readiness Lead
AI Data Readiness Lead

Deepgram • United States

On-site
USD 120,000 - 180,000
Principal AI Engineer
Principal AI Engineer

Logic Hire Solutions LTD • United States

Hybrid
USD 180,000 - 260,000
Senior Director of Machine Learning
Senior Director of Machine Learning

Hims & Hers • United States

Remote
USD 200,000 - 260,000
Generous PTO
Healthcare coverage
401(k) matching
+5
Analytics Engineer
Analytics Engineer

Lever, Inc. • Orem (UT)

On-site
USD 100,000 - 140,000