Get more replies from employers
Send a job-specific resume in minutes.
ScienceLogic is seeking a data scientist to advance our suite of locally-hosted language models and their evaluation framework. You will design evaluation harnesses, measure grounding and reliability, and drive cross‑functional improvements in an enterprise setting with security and compliance constraints.
You will collaborate with data scientists, ML/inference engineers, frontend, and product teams to turn interaction data into actionable insights, forecasting, anomaly detection, and robust
ScienceLogic is redefining IT operations for the modern enterprise. Our AIOps platform empowers organizations to achieve Autonomic IT — where systems are self‑healing, self‑optimizing, and seamlessly aligned with business outcomes. We help enterprises and service providers gain unified visibility across hybrid and multi‑cloud environments, automate workflows, and unlock performance at scale.
We’re accelerating digital transformation through the power of automation, AI, and analytics — giving IT and business leaders the tools to deliver superior customer experiences, drive efficiency, and innovate with confidence.
We're looking for a strong Data Scientist to join our growing Data Science team. We run a suite of small, locally‑hosted language models in production — not a single frontier API. That deliberate architecture defines this role: each model is more constrained than a giant hosted one, so product quality comes from how well we evaluate, route, prompt, ground, and orchestrate the models we have. Your job is to get the best possible outcomes out of that suite.
This is not classical predictive modeling. The object of measurement is the LLM system itself — its answers, retrieval, multi‑step agent behavior, and reliability under adversarial and edge‑case conditions. You'll define what "good" means for a non‑deterministic system running on bounded local models, build the evaluation infrastructure that catches regressions, and turn interaction data into the analysis that tells engineering and product where to invest.
You'll also build production prediction and trend capability — forecasting, anomaly detection, and early‑warning signals over operational telemetry — that feeds directly into that system. You'll work across data scientists, ML/inference engineers, frontend, and product in an enterprise environment with real security and compliance constraints. If you think in eval suites, failure modes, and groundedness — and you're energized by squeezing reliable, high‑quality behavior out of small models under real resource budgets — this is the role.
Key Responsibilities
Design and own evaluation harnesses for LLM and agentic outputs — golden sets, regression suites, and rubric‑based scoring.
Build and calibrate LLM‑as‑judge pipelines; validate judges against human labels and control for their bias and variance.
Define and track response‑quality metrics: faithfulness/groundedness, hallucination rate, answer relevance and completeness, instruction‑following, and persona adherence.
Curate, version, and grow evaluation datasets as the product and its surfaces evolve.
Benchmark the models in the suite against each other to decide which model handles which task, and quantify the quality cost of running smaller, local models versus larger alternatives.
Adversarial & Robustness Testing
Red‑team the system: prompt injection, jailbreaks, tool‑misuse, and edge‑case discovery.
Design chaos and stress tests that probe model and agent reliability under degraded or hostile conditions.
Characterize failure modes and feed them back into guardrails and regression coverage.
Retrieval & Agentic Trajectory Analysis
Evaluate retrieval quality over the document corpus — recall@k, MRR/nDCG, context precision and recall — and run experiments on chunking, indexing, and hybrid retrieval strategies.
Analyze multi‑step agent trajectories: tool‑call correctness, trajectory efficiency, replayable‑state inspection, and guardrail‑breach behavior.
Assess intent classification and routing quality as measurable components, not black boxes.
Behavioral Regression & Drift
Build standing evaluation that catches quality and behavioral regressions when a model in the suite is swapped, upgraded, or re‑quantized, or when prompts and pipelines change.
Monitor output‑distribution and quality drift in production; distinguish genuine regressions from noise on stochastic outputs.
Recommend and validate fixes through the levers available with local models — prompt changes, retrieval and grounding adjustments, routing changes, or model selection.
Predictive & Trend Modeling
Build, ship, and own production models that forecast and surface trends from operational telemetry — capacity and resource forecasting, anomaly prediction, and early‑warning signals on metrics and logs.
Take these from prototype to production and keep them healthy: deployment, monitoring, recalibration, and retraining as data and behaviour shift.
Define accuracy and lead‑time metrics that matter operationally — precision/recall on predicted incidents, forecast error, how far ahead a signal fires — not just offline scores.
Wire predictive signals into the LLM and agentic layer so forecasts and trends feed reasoning, advisories, and operator‑facing recommendations.
Domain & Value Analytics