Data Scientist

ScienceLogic

United States

On-site

USD 120,000 - 180,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ScienceLogic is seeking a data scientist to advance our suite of locally-hosted language models and their evaluation framework. You will design evaluation harnesses, measure grounding and reliability, and drive cross‑functional improvements in an enterprise setting with security and compliance constraints.

You will collaborate with data scientists, ML/inference engineers, frontend, and product teams to turn interaction data into actionable insights, forecasting, anomaly detection, and robust

Qualifications

  • Experience designing evaluation frameworks for ML systems.
  • Strong background in NLP and LLM evaluation metrics.
  • Ability to work with cross-functional teams in an enterprise.
  • Experience with deployment and monitoring of ML models.

Responsibilities

  • Design and own evaluation harnesses for LLM and agent outputs.
  • Build LLM-as-judge pipelines; validate judges against human labels and control for bias.
  • Define and track response-quality metrics: faithfulness, groundedness, relevance, and instruction-following.
  • Curate, version, and grow evaluation datasets as the product surfaces evolve.
  • Benchmark the models to decide task allocations and cost implications.
  • Red-team the system: prompt injection, jailbreaks, edge-case discovery.
  • Design chaos and stress tests for reliability under adverse conditions.
  • Characterize failure modes and feed them back into guardrails and regression coverage.
  • Evaluate retrieval quality and agent trajectories: recall@k, MRR, context precision/recall.
  • Assess intent classification and routing quality as measurable components.
  • Build production signals for forecasting, anomaly detection, and early-warning indicators.
  • Deploy, monitor, recalibrate models as data shifts, define actionable metrics.

Skills

LLM evaluation
Experimentation design
Data analysis
Python

Education

Master's degree in CS/Statistics or related field

Tools

PyTorch
scikit-learn

Job description

About ScienceLogic…

ScienceLogic is redefining IT operations for the modern enterprise. Our AIOps platform empowers organizations to achieve Autonomic IT — where systems are self‑healing, self‑optimizing, and seamlessly aligned with business outcomes. We help enterprises and service providers gain unified visibility across hybrid and multi‑cloud environments, automate workflows, and unlock performance at scale.
We’re accelerating digital transformation through the power of automation, AI, and analytics — giving IT and business leaders the tools to deliver superior customer experiences, drive efficiency, and innovate with confidence.

We're looking for a strong Data Scientist to join our growing Data Science team. We run a suite of small, locally‑hosted language models in production — not a single frontier API. That deliberate architecture defines this role: each model is more constrained than a giant hosted one, so product quality comes from how well we evaluate, route, prompt, ground, and orchestrate the models we have. Your job is to get the best possible outcomes out of that suite.

This is not classical predictive modeling. The object of measurement is the LLM system itself — its answers, retrieval, multi‑step agent behavior, and reliability under adversarial and edge‑case conditions. You'll define what "good" means for a non‑deterministic system running on bounded local models, build the evaluation infrastructure that catches regressions, and turn interaction data into the analysis that tells engineering and product where to invest.

You'll also build production prediction and trend capability — forecasting, anomaly detection, and early‑warning signals over operational telemetry — that feeds directly into that system. You'll work across data scientists, ML/inference engineers, frontend, and product in an enterprise environment with real security and compliance constraints. If you think in eval suites, failure modes, and groundedness — and you're energized by squeezing reliable, high‑quality behavior out of small models under real resource budgets — this is the role.

Key Responsibilities

  • Design and own evaluation harnesses for LLM and agentic outputs — golden sets, regression suites, and rubric‑based scoring.

  • Build and calibrate LLM‑as‑judge pipelines; validate judges against human labels and control for their bias and variance.

  • Define and track response‑quality metrics: faithfulness/groundedness, hallucination rate, answer relevance and completeness, instruction‑following, and persona adherence.

  • Curate, version, and grow evaluation datasets as the product and its surfaces evolve.

  • Benchmark the models in the suite against each other to decide which model handles which task, and quantify the quality cost of running smaller, local models versus larger alternatives.

Adversarial & Robustness Testing

  • Red‑team the system: prompt injection, jailbreaks, tool‑misuse, and edge‑case discovery.

  • Design chaos and stress tests that probe model and agent reliability under degraded or hostile conditions.

  • Characterize failure modes and feed them back into guardrails and regression coverage.

Retrieval & Agentic Trajectory Analysis

  • Evaluate retrieval quality over the document corpus — recall@k, MRR/nDCG, context precision and recall — and run experiments on chunking, indexing, and hybrid retrieval strategies.

  • Analyze multi‑step agent trajectories: tool‑call correctness, trajectory efficiency, replayable‑state inspection, and guardrail‑breach behavior.

  • Assess intent classification and routing quality as measurable components, not black boxes.

Behavioral Regression & Drift

  • Build standing evaluation that catches quality and behavioral regressions when a model in the suite is swapped, upgraded, or re‑quantized, or when prompts and pipelines change.

  • Monitor output‑distribution and quality drift in production; distinguish genuine regressions from noise on stochastic outputs.

  • Recommend and validate fixes through the levers available with local models — prompt changes, retrieval and grounding adjustments, routing changes, or model selection.

Predictive & Trend Modeling

  • Build, ship, and own production models that forecast and surface trends from operational telemetry — capacity and resource forecasting, anomaly prediction, and early‑warning signals on metrics and logs.

  • Take these from prototype to production and keep them healthy: deployment, monitoring, recalibration, and retraining as data and behaviour shift.

  • Define accuracy and lead‑time metrics that matter operationally — precision/recall on predicted incidents, forecast error, how far ahead a signal fires — not just offline scores.

  • Wire predictive signals into the LLM and agentic layer so forecasts and trends feed reasoning, advisories, and operator‑facing recommendations.

Domain & Value Analytics

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist, LLM Evaluation & AIOps
Data Scientist, LLM Evaluation & AIOps

Sciencelogic • United States

On-site
USD 120,000 - 155,000
401(k) plan with employer match
Flexible Paid Time Off (FTO)
Volunteer Time Off (VTO) - two days a 
+2
Senior Consultant, AI/ML Engineer
Senior Consultant, AI/ML Engineer

Hollstadt Consulting • Minnesota

On-site
USD 140,000 - 200,000
Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
LLM Systems Data Scientist — Evaluation, Robustness & Forecasting
LLM Systems Data Scientist — Evaluation, Robustness & Forecasting

ScienceLogic • United States

On-site
USD 120,000 - 180,000
Applied AI Engineer
Applied AI Engineer

SherlockTalent • Miami (FL)

Hybrid
USD 120,000 - 140,000
Solid Benefits
Referral bonus of $2,500
Data Scientist
Data Scientist

KANINI • Nashville (TN)

On-site
USD 120,000 - 180,000
Large Language Model Engineer - AI & Human Health Research
Large Language Model Engineer - AI & Human Health Research

Mount Sinai Health System • United States

On-site
USD 110,000 - 165,000
LLM Training & Model Development Engineer
LLM Training & Model Development Engineer

InOpTra Digital • United States

Remote
USD 90,000 - 120,000
Competitive salary
Opportunity for remote work
Health benefits
AI Engineer - GA
AI Engineer - GA

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
AI Engineer - FL
AI Engineer - FL

LawPro.ai • Town of Florida (NY)

On-site
USD 140,000 - 210,000