Sr. Evaluation Engineer

logicmonitor

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

LogicMonitor is seeking a Senior AI Engineer, Evaluations, to design and build evaluation systems for Edwin AI. You will create production-grade evaluation pipelines, golden datasets, automated graders, and regression frameworks for AI agents, retrieval systems, tool integrations, and complex investigation workflows.

You will define quality metrics, build offline and online evaluation pipelines in Python, integrate with CI/CD, and mentor others while monitoring AI quality and behavior drift in

Qualifications

  • 5+ years of experience in software engineering, ML, or a related field.

Responsibilities

  • Define quality metrics for incident diagnostics, root-cause analysis, alert correlation, grounding, tool use, safety, and operational usefulness.

Skills

Python
AI evaluation
Experimentation
Quality frameworks
LLM evaluation

Tools

LangSmith
Arize Phoenix
Braintrust
DeepEval
Ragas
TruLens
OpenAI Evals
MLflow

Job description

About Us

We love going to work and think you should too. Our team is dedicated to trust, customer obsession, agility, and striving to be better everyday. These values serve as the foundation of our culture, guiding our actions and driving us towards excellence. We foster a culture of performance and recognition, allowing us to transform growth as we enable our employees to do the best work of their careers.

This role is open to candidates based in or near San Francisco, CA. At LogicMonitor, we hire within our Centers of Energy-vibrant locations where our teams connect, collaborate, and innovate.

To learn more about life at LogicMonitor, check out our Careers Page.

What You'll Do

LogicMonitor® is the AI-first hybrid observability platform powering the next generation of digital infrastructure. LogicMonitor delivers complete visibility and actionable intelligence across on-premises, cloud, and edge environments. By anticipating issues before they strike, optimizing resources in real time, and enabling faster, smarter decisions, LogicMonitor helps IT and business leaders protect margins, accelerate innovation, and deliver exceptional digital experiences without compromise.

Our customers love LogicMonitor's ability to bring cloud and traditional IT together into one view, as seen in minimal churn rates, expansion business, and exciting new customer references. In fact, LogicMonitor has received the highest Net Promoter Score of any IT Infrastructure Management provider. LogicMonitor also boasts high employee satisfaction. We have been certified as a Great Place To Work®, and named one of BuiltIn's Best Places to Work for the seventh year in a row!

Edwin AI is LogicMonitor's AI-powered observability and incident intelligence platform. It helps enterprise operations teams investigate incidents, identify root causes, recommend remediation, and automate operational workflows.

As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide how Edwin AI is developed, tested, and released. You will create production-grade evaluation pipelines, golden datasets, automated graders, and regression frameworks for AI agents, retrieval systems, tool integrations, and complex investigation workflows.

Here’s a closer look at this key role:
  • Define quality metrics for incident diagnostics, root-cause analysis, alert correlation, grounding, tool use, safety, and operational usefulness.
  • Build offline and online evaluation pipelines in Python and integrate them with CI/CD, experimentation, model selection, prompt iteration, and release gating.
  • Lead the creation and maintenance of golden datasets and regression suites using alerts, events, metrics, logs, traces, topology, configuration data, incident timelines, change records, ITSM workflows, and historical investigation outcomes.
  • Build representative, customer-specific scenarios covering different technologies, failure modes, operational patterns, and environmental constraints.
  • Use human-authored and AI-assisted methods to generate regression, edge, adversarial, rare, ambiguous, and incomplete-context test cases.
  • Treat evaluation datasets and test suites as first-class components that evolve alongside Edwin AI.
  • Design step-level and trajectory-level evaluations for multi-step and multi-agent workflows.
  • Assess both final outcomes and intermediate behavior, including planning, reasoning consistency, retrieval, evidence use, tool selection, tool parameters, state transitions, escalation decisions, and human-in-the-loop approvals.
  • Identify whether failures originate from models, prompts, retrieval, data quality, tools, agent logic, orchestration, and infrastructure.
  • Evaluate capabilities including incident investigation, on-call assistance, operational question answering, change-impact analysis, remediation recommendations, automated resolution, infrastructure operations, and ITSM and observability integrations.
  • Design and calibrate LLM-based graders against expert human judgment.
  • Monitor AI quality and behavioral drift in production, and convert failures and customer feedback into new tests and safeguards.
  • Establish evaluation-driven development practices and mentor other engineers.
What You’ll Need
  • 5+ years of experience in software engineering, machine learning, applied AI, or a related field.
  • Strong Python engineering skills and experience building production systems.
  • Hands-on experience with AI evaluation, experimentation, testing, and quality frameworks.
  • Experience using multiple LLM and agent evaluation frameworks, such as LangSmith, Arize Phoenix, Braintrust, DeepEval, Ragas, TruLens, OpenAI Evals, MLflow, or comparable platforms.
  • Ability to select, customize, and integrate evaluation frameworks for offline testing, online monitoring, regression analysis, experimentation, model and prompt comparison, and release gating.
  • Strong understanding of LLMs, agents, retrieval-augmented generation, prompt engineering, tool
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Town of Poland (NY)

On-site
USD 140,000 - 200,000
Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Spain (TX)

On-site
EUR 70,000 - 100,000
Senior AI Evaluation Engineer: Pipelines & Metrics
Senior AI Evaluation Engineer: Pipelines & Metrics

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Evaluation Scientist
AI Evaluation Scientist

Steampunk, Inc. • McLean (VA)

On-site
USD 140,000 - 210,000
Observability and Evaluation Engineer
Observability and Evaluation Engineer

Mphasis • Charlotte (NC)

On-site
USD 120,000 - 180,000
AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
Senior Software Engineer - Model Evaluation & AI Systems
Senior Software Engineer - Model Evaluation & AI Systems

Worky • California (MO)

On-site
USD 180,000 - 230,000
LLM Evaluation Engineer
LLM Evaluation Engineer

ThirdLaw | Runtime AI Safety • United States

Hybrid
USD 120,000 - 160,000
Market cash compensation
Above-market equity
Generous benefits
QA / ML Tester — Evaluation Framework
QA / ML Tester — Evaluation Framework

Intellias • Town of Poland (NY)

On-site
USD 120,000 - 160,000
Python Engineer — Evaluator Library
Python Engineer — Evaluator Library

Intellias • Town of Poland (NY)

On-site
USD 120,000 - 180,000