LLM Systems Data Scientist — Evaluation, Robustness & Forecasting

ScienceLogic

United States

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

ScienceLogic is seeking a data scientist to advance our suite of locally-hosted language models and their evaluation framework. You will design evaluation harnesses, measure grounding and reliability, and drive cross‑functional improvements in an enterprise setting with security and compliance constraints.

You will collaborate with data scientists, ML/inference engineers, frontend, and product teams to turn interaction data into actionable insights, forecasting, anomaly detection, and robust

Qualifications

  • Experience designing evaluation frameworks for ML systems.
  • Strong background in NLP and LLM evaluation metrics.
  • Ability to work with cross-functional teams in an enterprise.
  • Experience with deployment and monitoring of ML models.

Responsibilities

  • Design and own evaluation harnesses for LLM and agent outputs.
  • Build LLM-as-judge pipelines; validate judges against human labels and control for bias.
  • Define and track response-quality metrics: faithfulness, groundedness, relevance, and instruction-following.
  • Curate, version, and grow evaluation datasets as the product surfaces evolve.
  • Benchmark the models to decide task allocations and cost implications.
  • Red-team the system: prompt injection, jailbreaks, edge-case discovery.
  • Design chaos and stress tests for reliability under adverse conditions.
  • Characterize failure modes and feed them back into guardrails and regression coverage.
  • Evaluate retrieval quality and agent trajectories: recall@k, MRR, context precision/recall.
  • Assess intent classification and routing quality as measurable components.
  • Build production signals for forecasting, anomaly detection, and early-warning indicators.
  • Deploy, monitor, recalibrate models as data shifts, define actionable metrics.

Skills

LLM evaluation
Experimentation design
Data analysis
Python

Education

Master's degree in CS/Statistics or related field

Tools

PyTorch
scikit-learn

Job description

ScienceLogic is seeking a data scientist to advance our suite of locally-hosted language models and their evaluation framework. You will design evaluation harnesses, measure grounding and reliability, and drive cross‑functional improvements in an enterprise setting with security and compliance constraints.

You will collaborate with data scientists, ML/inference engineers, frontend, and product teams to turn interaction data into actionable insights, forecasting, anomaly detection, and robust

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Scientist, LLM Evaluation & AIOps
Data Scientist, LLM Evaluation & AIOps

Sciencelogic • United States

On-site
USD 120,000 - 155,000
401(k) plan with employer match
Flexible Paid Time Off (FTO)
Volunteer Time Off (VTO) - two days a 
+2
Data Scientist
Data Scientist

ScienceLogic • United States

On-site
USD 120,000 - 180,000
LLM Data Scientist: End-to-End Training & Evaluation
LLM Data Scientist: End-to-End Training & Evaluation

Propio Language Services • Overland Park (KS)

On-site
USD 120,000 - 180,000
Multilingual LLM Data Engineer & Evaluation Scientist
Multilingual LLM Data Engineer & Evaluation Scientist

Propio • Overland Park (KS)

Hybrid
USD 120,000 - 180,000
Staff Data Scientist — LLM Data Analysis & Evaluation
Staff Data Scientist — LLM Data Analysis & Evaluation

Cohere • New York (NY), San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health benefits
RRSP matching
+6
Staff Data Scientist, LLM Data Quality & Evaluation
Staff Data Scientist, LLM Data Quality & Evaluation

Cohere • San Francisco (CA)

Remote
USD 120,000 - 160,000
Weekly lunch stipend
Full health and dental benefits
401K and pension scheme
+3
NLP Data Scientist ML & LLMs for Impact
NLP Data Scientist ML & LLMs for Impact

j5holdingscorporation • Chantilly (VA)

On-site
USD 120,000 - 160,000
Health coverage
Tuition reimbursement
PTO
+2
Senior AI Engineer — LLM Evaluation & Production Systems
Senior AI Engineer — LLM Evaluation & Production Systems

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
Senior Data Scientist: LLM Eval & BI for Trials
Senior Data Scientist: LLM Eval & BI for Trials

Paradigm Health, Inc. • Town of Columbus (NY)

Hybrid
USD 120,000 - 180,000
LLMOps Data Analyst: Observability & Evaluation Leader
LLMOps Data Analyst: Observability & Evaluation Leader

ADP, Inc. • El Segundo (CA)

On-site
USD 110,000 - 120,000