Observability and Evaluation Engineer

Mphasis

Charlotte (NC)

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mphasis is seeking an Observability and Evaluation Engineer in Charlotte, NC to design and implement observability pipelines for AI/LLM workflows, and build robust evaluation and testing frameworks to ensure model quality and reliability.

You will instrument applications, develop offline/online evaluation suites, and work on performance and cost optimization across distributed microservices and AI orchestration layers.

Qualifications

  • Strong proficiency in Python or Go for building custom evaluation scripts and telemetry hooks.
  • Experience with AI observability platforms and telemetry stacks for tracking LLM/agent workflows.
  • Knowledge of adversarial prompts and testing for LLM/system validation.

Responsibilities

  • Build Observability Pipelines: instrument applications and LLM/agent pipelines using telemetry tools to capture traces, logs, and latency metrics.
  • Design Evaluation Frameworks: create offline and online evaluation suites to benchmark accuracy, groundedness, toxicity, tool-use correctness, and reasoning-chain validity.
  • Implement Regression & Drift Testing: build automated test harnesses and continuous evaluation gates to detect model degradation, data drift, or output anomalies before releases reach production.
  • Root-Cause Analysis: investigate execution traces and multi-turn interaction failures to diagnose erratic system behaviors, API misparameters, or bottlenecks.
  • Optimize Performance & Cost: monitor and balance operational telemetry relating to token consumption, execution speed, and infrastructure costs.

Skills

Python
Go
Testing methodologies
System design

Tools

OpenTelemetry
Prometheus
Grafana
Arize Phoenix
LangChain/LangSmith
Galileo
Braintrust

Job description

Core Responsibilities
  • Build Observability Pipelines: Instrument applications and LLM/agent pipelines using telemetry tools (like OpenTelemetry, Arize, Galileo, or LangSmith) to capture execution traces, run logs, and latency metrics. [1]
  • Design Evaluation Frameworks: Create offline and online evaluation suites to benchmark accuracy, groundedness, toxicity, tool-use correctness, and reasoning-chain validity. []
  • Implement Regression & Drift Testing: Build automated test harnesses and continuous evaluation gates to detect model degradation, data drift, or output anomalies before releases reach production. [1]
  • Root-Cause Analysis: Investigate execution traces and multi-turn interaction failures to diagnose erratic system behaviors, API misparameters, or bottlenecks. []
  • Optimize Performance & Cost: Monitor and balance operational telemetry relating to token consumption, execution speed, and infrastructure costs. [1]
Key Skills & Requirements
  • Programming: Strong proficiency in Python or Go for building custom evaluation scripts and hooking into telemetry SDKs.
  • Tools & Stacks: Experience with OpenTelemetry, Prometheus, Grafana, or dedicated AI observability platforms (e.g., Arize Phoenix, LangChain/LangSmith, Galileo, Braintrust).
  • Testing Methodologies: Designing adversarial prompts, red-teaming protocols, and golden datasets for LLM/system validation.
  • System Design: Understanding of distributed microservices or LLM orchestration layers (LangChain, LlamaIndex, AutoGen). [1, 2, 3]

An Observability and Evaluation Engineer (often specialized in AI/LLM systems) designs the tracking, tracing, and testing frameworks that monitor production behavior, measure output quality, and catch regressions in complex software or agentic AI workflows. [1, 2, 3, 4]

If you are tailoring this for a specific application, would you like me to focus this job description more heavily on traditional cloud/microservices observability or Generative AI / LLM agent evaluation?

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Evaluation Engineer
Sr. Evaluation Engineer

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Town of Poland (NY)

On-site
USD 140,000 - 200,000
Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Spain (TX)

On-site
EUR 70,000 - 100,000
Senior AI Engineer
Senior AI Engineer

arosplatforms | AI Consulting & Services • North Township (IN)

On-site
USD 120,000 - 170,000
QA / ML Tester — Evaluation Framework
QA / ML Tester — Evaluation Framework

Intellias • Town of Poland (NY)

On-site
USD 120,000 - 160,000
AI Engineer - GA
AI Engineer - GA

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Applied AI Engineer
Applied AI Engineer

SherlockTalent • Miami (FL)

Hybrid
USD 120,000 - 140,000
Solid Benefits
Referral bonus of $2,500
AI Engineer - OH
AI Engineer - OH

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
AI Engineer - NC
AI Engineer - NC

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity