LLM Systems Data Scientist — Evaluation, Robustness & Forecasting

ScienceLogic

United States

On-site

USD 120,000 - 180,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ScienceLogic is seeking a data scientist to advance our suite of locally-hosted language models and their evaluation framework. You will design evaluation harnesses, measure grounding and reliability, and drive cross‑functional improvements in an enterprise setting with security and compliance constraints.

You will collaborate with data scientists, ML/inference engineers, frontend, and product teams to turn interaction data into actionable insights, forecasting, anomaly detection, and robust

Qualifications

  • Experience designing evaluation frameworks for ML systems.
  • Strong background in NLP and LLM evaluation metrics.
  • Ability to work with cross-functional teams in an enterprise.
  • Experience with deployment and monitoring of ML models.

Responsibilities

  • Design and own evaluation harnesses for LLM and agent outputs.
  • Build LLM-as-judge pipelines; validate judges against human labels and control for bias.
  • Define and track response-quality metrics: faithfulness, groundedness, relevance, and instruction-following.
  • Curate, version, and grow evaluation datasets as the product surfaces evolve.
  • Benchmark the models to decide task allocations and cost implications.
  • Red-team the system: prompt injection, jailbreaks, edge-case discovery.
  • Design chaos and stress tests for reliability under adverse conditions.
  • Characterize failure modes and feed them back into guardrails and regression coverage.
  • Evaluate retrieval quality and agent trajectories: recall@k, MRR, context precision/recall.
  • Assess intent classification and routing quality as measurable components.
  • Build production signals for forecasting, anomaly detection, and early-warning indicators.
  • Deploy, monitor, recalibrate models as data shifts, define actionable metrics.

Skills

LLM evaluation
Experimentation design
Data analysis
Python

Education

Master's degree in CS/Statistics or related field

Tools

PyTorch
scikit-learn

Job description

ScienceLogic is seeking a data scientist to advance our suite of locally-hosted language models and their evaluation framework. You will design evaluation harnesses, measure grounding and reliability, and drive cross‑functional improvements in an enterprise setting with security and compliance constraints.

You will collaborate with data scientists, ML/inference engineers, frontend, and product teams to turn interaction data into actionable insights, forecasting, anomaly detection, and robust

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist, LLM Evaluation & AIOps
Data Scientist, LLM Evaluation & AIOps

Sciencelogic • United States

On-site
USD 120,000 - 155,000
401(k) plan with employer match
Flexible Paid Time Off (FTO)
Volunteer Time Off (VTO) - two days a 
+2
Data Scientist
Data Scientist

ScienceLogic • United States

On-site
USD 120,000 - 180,000
Staff Data Scientist — LLM Data Analysis & Evaluation
Staff Data Scientist — LLM Data Analysis & Evaluation

Cohere • New York (NY), San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health benefits
RRSP matching
+6
Staff Data Scientist, LLM Data Quality & Evaluation
Staff Data Scientist, LLM Data Quality & Evaluation

Cohere • San Francisco (CA)

Remote
USD 120,000 - 160,000
Weekly lunch stipend
Full health and dental benefits
401K and pension scheme
+3
LLM Evaluation & Post-Training Scientist
LLM Evaluation & Post-Training Scientist

Innodata Inc. • United States

On-site
USD 175,000 - 225,000
Senior LLM Evaluation Scientist
Senior LLM Evaluation Scientist

cohere • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health benefits
RRSP matching
+5
Senior AI Engineer — LLM Evaluation & Production Systems
Senior AI Engineer — LLM Evaluation & Production Systems

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
LLM Evaluation Engineer — Quality & Automation
LLM Evaluation Engineer — Quality & Automation

Grid Dynamics • United States

On-site
USD 140,000 - 170,000
Flexible schedule
Medical insurance
Vision and dental
+3
LLM Systems Architect & Advisor
LLM Systems Architect & Advisor

Peraton • Idaho

On-site
USD 104,000 - 166,000
Senior Applied AI Scientist - LLM Evaluation
Senior Applied AI Scientist - LLM Evaluation

Dadi Inc (acquired by Ro) • New York (NY)

On-site
USD 182,000 - 220,000
Competitive equity
Benefits package
Health benefits