Agentic QA Engineer Generative AI & Agentic Systems (Agent, Multi Agent Testing)
Location: Dallas, TX - Onsite
Duration: Long term
Summary
We are seeking a hands-on AI Engineer to design and execute end-to-end testing strategies for agentic AI solutions, including multi-agent systems in production-grade environments. This role partners with the Agentic Operations Team to ensure resiliency, reliability, accuracy, latency, orchestration correctness, and scale.
Key Responsibilities
- Agentic & Multi Agent Testing
- Reliability, Resiliency, and Latency
- Accuracy & Macro-Level Validations
- Scale & Orchestration
- Dev Prod Readiness
- Define and own the QA strategy for agentic/multi-agent AI systems across dev, staging, and prod.
- Mentor a team of QA engineers; establish testing standards, coding guidelines for test harnesses, and review practices.
- Partner with Agentic Operations, Data Science, MLOps, and Platform teams to embed QA in the SDLC and incident response.
- Design tests for agent orchestration, tool calling, planner-executor loops, and
- Establish latency SLOs and measure end-to-end response times across orchestration layers (LLM calls, tool invocations, queues).
- Ensure reliability through soak tests, canary verifications, and automated rollbacks.
- Define ground-truth and reference pipelines for task accuracy (exact match, semantic similarity, factuality checks).
- Build macro validation frameworks that validate task outcomes across multi-step agent workflows (e.g., complex data pipelines, content generation + verification agent loops).
- Instrument guardrail validations (toxicity, PII, hallucination, policy compliance).
Required Qualifications
- 7+ years in Software QA/Testing, with 2+ years in AI/ML or LLM-based systems; hands-on experience testing agentic/multi-agent architectures.
- Strong programming skills in Python experience building test harnesses, simulators, and fixtures.
- Experience with LLM evaluation (exact/soft match, BLEU/ROUGE, BERTScore, semantic similarity via embeddings), guardrails, and prompt testing.
- Expertise in distributed systems testing latency profiling, resiliency patterns (circuit breakers, retries), chaos engineering, and message queues.
- Familiarity with orchestration frameworks (LangChain, LangGraph, LlamaIndex, DSPy, OpenAI Assistants/Actions, Azure OpenAI orchestration, or similar).
- Proficiency with CI/CD (GitHub Actions/Azure DevOps), observability (OpenTelemetry, Prometheus/Grafana, Datadog), and feature flags/canaries.
Preferred Qualifications
- Experience with multi-agent simulators, agent graph testing, and tooling latency emulation.
- Knowledge of MLOps (model versioning, datasets, evaluation pipelines) and A/B experimentation for LLMs.
- Background in cloud (AWS), serverless, containerization, and event-driven architectures.