AI Quality Engineer: Evaluation Pipelines & LLM Testing

Dealstitch LLC.

United States

On-site

USD 160,000 - 220,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Dealstitch LLC is seeking an AI Quality Engineer to own the evaluation system for production-grade agents. You will run and extend pipelines that validate agents in client engagements and our internal fleet, ensuring honesty as models and data evolve.

You will measure success with automated gates in CI/CD, score agents at step-level, and expand the toolkit for failures, safety, and observability across production deployments.

Qualifications

  • Strong Python programming and testing for non-deterministic systems.
  • Experience evaluating LLM/agent systems: agent graphs, eval frameworks, and LLM-as-a-judge methods.
  • Fluency with retrieval quality (RAG), observability and tracing for production agents.
  • A bias for measurement: define what "good" means and prove it with a harness.
  • Bonus: data-pipeline testing or exposure to private markets or enterprise data.

Responsibilities

  • Own the evaluation pipelines that gate every change we ship, automated in CI/CD.
  • Score agents at the step level across tool selection, planning, reasoning chains, and retrieval quality.
  • Extend the eval toolkit for new failure modes: LLM-as-a-judge, RAG faithfulness, hallucination detection, and red-teaming for safety.
  • Maintain golden datasets and tracing/observability for live agents.

Skills

Python
Testing non-deterministic systems
CI/CD
LLM eval frameworks
Observability/Tracing

Job description

Dealstitch LLC is seeking an AI Quality Engineer to own the evaluation system for production-grade agents. You will run and extend pipelines that validate agents in client engagements and our internal fleet, ensuring honesty as models and data evolve.

You will measure success with automated gates in CI/CD, score agents at step-level, and expand the toolkit for failures, safety, and observability across production deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer: Scale QA for LLMs
AI Evaluation Engineer: Scale QA for LLMs

Appnovation • Dallas (TX)

On-site
USD 95,000 - 140,000
Senior AI Evaluation Engineer: Pipelines & Metrics
Senior AI Evaluation Engineer: Pipelines & Metrics

logicmonitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Senior AI Quality Engineer — LLMs & Evaluation Systems
Senior AI Quality Engineer — LLMs & Evaluation Systems

Block • San Francisco (CA)

On-site
USD 190,000 - 230,000
Staff AI Systems Engineer - LLM, Evaluation, Equity
Staff AI Systems Engineer - LLM, Evaluation, Equity

Maven • San Jose (CA)

On-site
USD 180,000 - 240,000
Physical Health Benefits
Mental Health Benefits
Emotional Health Benefits
+5
AI Agent Engineer: Build LLM Pipelines & Graph-RAG
AI Agent Engineer: Build LLM Pipelines & Graph-RAG

Appnovation • New York (NY)

On-site
USD 120,000 - 170,000
Senior AI Analytics Engineer: Data Pipelines & LLM Quality
Senior AI Analytics Engineer: Data Pipelines & LLM Quality

Socket.dev • Mountain View (CA)

On-site
USD 150,000 - 260,000
Health insurance
Employee stock ownership
Generous vacation and personal days
+2
AI Agent Developer QA Specialization
AI Agent Developer QA Specialization

Accord Technologies Inc • New Jersey

On-site
USD 90,000 - 120,000
Senior AI Engineer: Agentic LLM Pipelines & Systems
Senior AI Engineer: Agentic LLM Pipelines & Systems

Strativ Group • Boston (MA)

On-site
USD 140,000 - 190,000
AI Agent Engineer: Build LLM Tools & RAG Pipelines
AI Agent Engineer: Build LLM Tools & RAG Pipelines

Appnovation Technologies • New York (NY)

On-site
USD 140,000 - 190,000
AI Product Engineer - LLMs, Agents & Workflow Automation
AI Product Engineer - LLMs, Agents & Workflow Automation

Cohere • New York (NY)

On-site
USD 140,000 - 210,000
Weekly Lunch Stipend
Health Benefits
Dental Benefits
+15