Agentic QA Engineer/Generative AI & Agentic Systems | Dallas, TX – 5 Days Onsite | Must Skills: Python, AI/ML, LLMs, Agentic/Multi-Agent Testing, LLM Evaluation, Prompt Testing, LangChain/LangGraph/LlamaIndex, Distributed…

Keylent Inc

Dallas (TX)

On-site

USD 120,000 - 180,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Keylent Inc. is seeking a hands-on QA Engineer to design end-to-end testing strategies for agentic AI and multi-agent systems in production environments. You will partner with operations to ensure resiliency, reliability, and scale, establishing QA frameworks and reusable test assets across the SDLC.

The role requires leadership to mentor QA engineers and align with data, ops, and platform teams. The ideal candidate will have strong Python skills, experience with LLM evaluation, and a background

Qualifications

  • 7 years in Software QA/Testing with 2 years in AI/ML or LL/m-based systems.
  • Hands-on experience testing agentic/multi-agent architectures.

Responsibilities

  • Design and execute end-to-end testing strategies for agentic AI solutions.
  • Build reusable test artifacts and frameworks for multi-agent systems.
  • Lead QA across dev to prod with CI/CD integration.

Skills

Python
LLM evaluation
Distributed systems
CI/CD

Tools

LangChain
LangGraph
LlamaIndex
DSPy
Azure DevOps

Job description

JOB DESCRIPTION:

Agentic QA Engineer — Generative AI & Agentic Systems (Agent, Multi-Agent Testing)

Location: Dallas, TX

Duration: Long term

Mode of interview: One Virtual and One face to face

Mode of Job: 5 Days onsite

Summary

We are seeking a hands‑on AI Engineer to design and execute end‑to‑end testing strategies for agentic AI solutions, including multi‑agent systems in production‑grade environments. This role partners with the Agentic Operations Team to ensure resiliency, reliability, accuracy, latency, orchestration correctness, and scale. You will establish QA frameworks, build reusable test artifacts, drive macro‑level validations across complex workflows, and lead the QA function for Agentic AI from Dev to Prod.

Key Responsibilities
  • Agentic & Multi‑Agent Testing
  • Reliability, Resiliency, and Latency
  • Accuracy & Macro‑Level Validations
  • Scale & Orchestration
  • Dev Prod Readiness
  • Define and own the QA strategy for agentic/multi‑agent AI systems across dev, staging, and prod.
  • Mentor a team of QA engineers; establish testing standards, coding guidelines for test harnesses, and review practices.
  • Partner with Agentic Operations, Data Science, MLOps, and Platform teams to embed QA in the SDLC and incident response.
  • Design tests for agent orchestration, tool calling, planner‑executor loops, and inter‑agent coordination (e.g., task decomposition, handoff integrity, and convergence to goals).
  • Validate state management, context windows, memory/knowledge stores, and prompt/graph correctness under varying conditions.
  • Implement scenario fuzzing (e.g., adversarial inputs, prompt perturbations, tool latency spikes, degraded APIs).
  • Create resilience testing suites: chaos experiments, failover, retries/backoff, circuit‑breaking, and degraded mode behavior.
  • Establish latency SLOs and measure end‑to‑end response times across orchestration layers (LLM calls, tool invocations, queues).
  • Ensure reliability through soak tests, canary verifications, and automated rollbacks.Define ground‑truth and reference pipelines for task accuracy (exact match, semantic similarity, factuality checks).
  • Build macro validation frameworks that validate task outcomes across multi‑step agent workflows (e.g., complex data pipelines, content generation verification agent loops).
  • Instrument guardrail validations (toxicity, PII, hallucination, policy compliance).
  • Design load/stress tests for multi‑agent graphs under scale (concurrency, throughput, queue depth, backpressure).
  • Validate orchestrator correctness (DAG execution, retries, branching, timeouts, compensation paths).
  • Engineer reusable test artifacts (scenario configs, synthetic datasets, prompt libraries, agent graph fixtures, simulators).
  • Integrate tests into CI/CD (pre-merge gates, nightly, canary) and production monitoring with alerting tied to KPIs.
  • Define release criteria and run operational readiness (performance, security, compliance, cost/latency budgets).
Required Qualifications
  • 7 years in Software QA/Testing, with 2 years in AI/ML or LLM‑based systems; hands‑on experience testing agentic/multi‑agent architectures.
  • Strong programming skills in Python experience building test harnesses, simulators, and fixtures.
  • Experience with LLM evaluation (exact/soft match, BLEU/ROUGE, BERTScore, semantic similarity via embeddings), guardrails, and prompt testing.
  • Expertise in distributed systems testing latency profiling, resiliency patterns (circuit breakers, retries), chaos engineering, and message queues.
  • Familiarity with orchestration frameworks (LangChain, LangGraph, LlamaIndex, DSPy, OpenAI Assistants/Actions, Azure OpenAI orchestration, or similar).
  • Proficiency with CI/CD (GitHub Actions/Azure DevOps), observability (OpenTelemetry, PrometheGrafana, Datadog), and feature flags/canaries.
  • Solid understanding of privacy/security/compliance in AI systems (PII handling, content policies, model safety).
  • Excellent communication and leadership skills; proven ability to work cross‑functionally with Ops, Data, and Engineering.
Preferred Qualifications
  • Experience with multi‑agent simulators, agent graph testing, and tooling latency emulation.
  • Knowledge of MLOps (model versioning, datasets, evaluation pipelines) and A/B experimentation for LLMs.
  • Background in cloud (AWS), serverless, containerization, and event‑driven architectures.
  • Prior ownership of cost/latency/SLAs for AI workloads in production.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Software Engineer in Test – Agentic AI (Remote - US)
Lead Software Engineer in Test – Agentic AI (Remote - US)

Vibehackers • Northern (KY)

Hybrid
USD 150,000 - 190,000
Generous annual bonus opportunity
401(k) with employer match
Medical Insurance
+2
QA / Automation Engineer Agentic AI
QA / Automation Engineer Agentic AI

Compunnel, Inc. • Atlanta (GA), Northern (KY)

Hybrid
USD 110,000 - 160,000
QA / Automation Engineer— Agentic AI
QA / Automation Engineer— Agentic AI

Advanced Tech Placement • Roseland (NJ)

On-site
USD 90,000 - 130,000
Senior AI Engineer, Agentic Systems (Java, Python)
Senior AI Engineer, Agentic Systems (Java, Python)

Talentola • Plano (TX)

On-site
USD 180,000 - 240,000
AI QA lead
AI QA lead

Seneca Resources • New York (NY)

On-site
USD 180,000 - 240,000
Senior Quality Engineering / Agentic AI Lead
Senior Quality Engineering / Agentic AI Lead

Inabia Software & Consulting Inc. • Town of Charlotte (NY)

On-site
USD 140,000 - 180,000
AI Agentic Tester
AI Agentic Tester

Delan Associates, Inc • Denver (CO)

Hybrid
USD 90,000 - 120,000
Senior Agentic AI Engineer
Senior Agentic AI Engineer

ImagineX LLC • Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior Agentic AI Engineer
Senior Agentic AI Engineer

Build • Northern (KY)

Hybrid
USD 150,000 - 190,000
QA Engineer
QA Engineer

WebSenor Ltd • United States

On-site
USD 70,000 - 110,000