QA AI Automation Engineer

Dynasty Financial Partners, LLC

Saint Petersburg (FL)

On-site

USD 120,000 - 150,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Dynasty Financial Partners, LLC is building AI systems for advisors and clients within a regulated wealth management platform. This role focuses on evaluating non-deterministic systems, developing tooling to catch AI regressions, and ensuring code quality in agent-written software.

You will own AI evaluation, design multi-agent QA workflows, and scale automated gates across delivery teams, partnering with AI Labs and product to define coverage and acceptance criteria.

Qualifications

  • 5+ years in software quality, test automation, or SDET with ownership of automation architecture.
  • Production experience with LLM-powered systems: agent patterns, prompt engineering, tool calling, orchestration.
  • Hands-on with model APIs (OpenAI, Anthropic, Azure OpenAI) and AI-assisted development tooling.
  • Strong programming in Python, JavaScript/TypeScript, Java, Kotlin, or C#.
  • Experience building evaluation criteria for systems without a single correct answer.
  • Deep CI/CD fluency: pipeline integration, gating, monitoring, and logging.
  • API testing depth and validating third-party integrations.
  • Clear written communication to explain risks to product owners and engineers.

Responsibilities

  • Own and extend AI evaluation framework for AI chat and API surface.
  • Build multi-agent quality workflow in CI/CD: run, triage, defect writeups, fixes, re-run.
  • Define agent roles, guardrails, and auto-apply vs human review criteria.
  • Make cost-aware decisions on agent vs deterministic automation.
  • Report autonomous resolution rate, false positives, and engineer hours saved.
  • Scale code quality with automated gates for agent-authored code.
  • Partner with product, engineering, and AI Labs on coverage and acceptance.
  • Contribute to release-readiness and go/no-go decisions with evidence.
  • Maintain and modernize automation estate; mentor engineers.

Skills

Automation architecture
LLM experience
CI/CD fluency
API testing
Python
JavaScript/TypeScript
Communication

Tools

Langfuse
Promptfoo
OpenTelemetry
New Relic
n8n
Azure

Job description

Description

About the Role

We are building AI systems that advisors and their clients rely on inside a regulated wealth management platform — an internal AI assistant, an external-facing API layer on top of it, and a growing set of agentic workflows. The bar for correctness, grounding, and auditability is higher here than it is for most AI products.

This role exists because traditional QA does not cover that surface. We need an engineer who can evaluate non‑deterministic systems, build the tooling that catches AI regressions before they reach an advisor, and hold the line on code quality now that a meaningful share of our code is written by autonomous coding agents. This is a hands‑on engineering role, not a test‑execution role.

What You’ll Do
1. Own and extend our AI evaluation framework for AI chat
  • Build and maintain evaluation suite for our AI assistant and its API surface — correctness, grounding/citation fidelity, retrieval quality, refusal and escalation behavior, tone, latency, and cost per interaction.
  • Curate and version golden datasets and adversarial test sets drawn from real advisor workflows; keep them representative as the product changes.
  • Design LLM‑as‑judge and rubric‑based scoring where deterministic assertions do not apply and validate the judges themselves against human‑labeled sets.
  • Wire evals into CI so prompt, model, retrieval, and tool changes are gated by measured regression — not by vibes.
  • Instrument production traces and close the loop from live failures back into the eval suite.
2. Build the multi‑agent AI quality step that finds defects automatically
  • Design and operate a multi‑agent workflow in our CI/CD pipeline that runs on commit: executes the suite, triages failures, writes up defects with reproduction detail, proposes or applies fixes, and re‑runs to verify.
  • Define agent roles, hand‑off logic, guardrails, and the quality checks that decide when an agent's output is trustworthy enough to auto‑apply versus route to a human.
  • Make cost‑aware calls on where an LLM us needed and where deterministic automation is simply the better tool.
  • Report on the loop's real impact — autonomous resolution rate, false‑positive rate, engineer hours returned.
3. Scale code quality in the age of autonomous agent coding
  • Define what “good” means for agent‑authored code and build the automated gates that enforce it: coverage and mutation testing, static analysis, security and dependency scanning, architectural conformance, and review checklists tuned to how agents fail.
  • Identify the failure modes specific to agent‑generated code — plausible‑but‑wrong logic, silent scope creep, duplicated abstractions, missing edge‑case handling, tests written to pass rather than to verify — and build detection for them.
  • Partner with architects and delivery teams to keep velocity from outrunning quality as agent‑assisted development scales across the org.
Across all three
  • Partner with product, engineering, and AI Labs to define coverage and acceptance criteria for AI features before they are built.
  • Contribute to release‑readiness and go/no‑go decisions with evidence.
  • Maintain and modernize our existing automation estate (API, UI, integration) and mentor engineers on

Requirements

Required
  • 5+ years in software quality, test automation, or SDET work, with real ownership of automation architecture — not just test authoring.
  • Demonstrated production experience with LLM‑powered systems: agent patterns, prompt engineering, tool/function calling, and orchestration.
  • Hands‑on with major model APIs (OpenAI, Anthropic, Azure OpenAI) and AI‑assisted development tooling (Claude Code, Copilot, Codex or equivalent).
  • Strong programming in at least one of Python, TypeScript/JavaScript, Java, Kotlin, or C#, and comfort reading across the others.
  • Experience building and interpreting evaluation criteria for systems without a single correct answer.
  • Deep CI/CD fluency — pipeline integration, gating, monitoring, and logging.
  • API testing depth and experience validating third‑party integrations.
  • Clear written communication; you can explain a quality risk to a product owner and a root cause to an engineer in the same day.
Preferred
  • Eval and observability tooling: Langfuse, Promptfoo, OpenTelemetry, New Relic, or equivalents.
  • RAG fundamentals — embeddings, chunking strategy, vector search, retrieval evaluation.
  • Workflow orchestration tooling (n8n or similar).
  • Azure cloud services; .NET ecosystem exposure.
  • BDD/Selenium or comparable UI automation at scale.
  • Financial services, wealth management, or another regulated domain — or a clear appetite for what data governance means in one.
  • Experience mentoring or leading distributed QA engineers.
First 90 Days
  • Days 1–30: Learn the assistant, the API layer, and the current eval and automation estate. Ship a baseline eval suite for one high‑traffic chat workflow with results visible in CI.
  • Days 31–60: Stand up the first multi‑agent defect‑detection step on a single service; establish the trust threshold for auto‑applied fixes and publish the metrics.
  • Days 61–90: Propose and begin rolling out the agent‑authored code quality standard across delivery teams, with automated gates.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

QA AI Automation Engineer
QA AI Automation Engineer

Dynasty Financial Partners • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
QA AI Automation Engineer
QA AI Automation Engineer

Socket.dev • Saint Petersburg (FL)

On-site
USD 120,000 - 180,000
Lead Quality Automation Engineer – AI Platform
Lead Quality Automation Engineer – AI Platform

InSite • Washington

On-site
USD 120,000 - 170,000
Lead AI Quality Engineer
Lead AI Quality Engineer

Phase2 • Town of Montana (WI)

Hybrid
USD 150,000 - 190,000
Senior Applied AI Engineer
Senior Applied AI Engineer

Level • Austin (CO)

On-site
USD 180,000 - 240,000
Relocation assistance
QA Engineer (AI Systems)
QA Engineer (AI Systems)

Engg • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior Applied AI Engineer
Senior Applied AI Engineer

Level • Austin (TX)

On-site
USD 120,000 - 150,000
AI Engineer
AI Engineer

Valsoft Corporation • Northern (KY)

Hybrid
USD 120,000 - 180,000
AI Engineer
AI Engineer

Valsoft Corporation • United States

On-site
USD 140,000 - 230,000
AI Engineer
AI Engineer

eMAX Health Systems • New York (NY)

On-site
USD 170,000 - 210,000