QA Engineer: AI-Native Quality, Equity & Impact

TestDevJobs

Boston, Northern (MA, KY)

Hybrid

USD 115,000 - 130,000

Full time

23 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Newton Research, a Boston-area AI-first QA startup, seeks a QA Engineer focused on evals for AI-native agents. You will own curated eval datasets, scoring rubrics, and regression tracking across model, prompt, skill, and agent releases to ensure safe, reliable agent behavior.

You will test changes before release, validate tool-call correctness, and cover end-to-end agent flows. The role emphasizes non-determinism handling, data-driven decisions, and collaboration with engineering on testability.

Qualifications

  • 4+ years in QA or test engineering on a complex web product.
  • Hands-on experience building evals for LLM or agent products (datasets, rubrics, LLM-as-judge, regression tracking), or a clear track record of getting there fast
  • Working knowledge of how agents work: prompting, skills, tool calling, context, RAG and orchestration, well enough to tell where a failure originates
  • An automation-first instinct: you can show what you removed from a manual process
  • Python (or similar) for eval tooling and data checks; Playwright or similar for end-to-end tests
  • Statistical literacy: pass rates, variance and sample size when outputs are not deterministic
  • Daily user of AI coding and testing assistants, with the judgment to verify their output rather than trust it
  • Strong exploratory instincts and excellent written communication
  • Nice to have: adtech, martech or marketing analytics domain; SSO / SAML / OAuth flows; data-connector testing; eval or observability tooling

Responsibilities

  • Own the eval suite for Newton's agents: curated datasets of inputs and expected behaviors, rubric and LLM-as-judge scoring, regression tracking across model, prompt, skill and agent releases; validate judges against human-labeled examples
  • Test skill and prompt changes before they ship: every edit to a skill, system prompt or tool definition runs against the relevant evals in CI, with a before/after comparison a reviewer can read in one glance; gate releases on the results
  • Cover the agent flows end to end: tool-call correctness, task completion, multi-turn coherence, blueprint creation from conversations, code generation, scheduled tasks; flag output that is wrong, empty or silently degraded
  • Handle non-determinism with rigor: repeated runs, pass-rate thresholds and simple statistics, so a flaky agent is a measured finding, not an anecdote
  • Turn production and customer signal into evals: mine logs, error tracking and customer reports so every escaped bad-behavior case becomes a permanent eval, ideally drafted by an agent and reviewed by you
  • Probe AI-specific risk: prompt injection, data leakage across users, projects and permissions, hallucinated or ungrounded numbers in analytics output, cost and latency regressions
  • Automate before you repeat: any check you do by hand twice becomes a Playwright test, an eval or an agent workflow; drive manual regression time down every sprint
  • Design agentic QA workflows in CI: agents that run suites, triage failures, draft defects and re-verify fixes, with guardrails, cost limits and escalation rules you define
  • Keep human-judgment work sharp: exploratory testing across roles, feature flags and environments, SSO and connector authorization, and the cross-feature bugs only a person thinks to look for
  • Write bug reports that are machine- and human-readable, and partner with engineering on testability (observability, seedable data, stable interfaces)

Skills

QA testing
Eval tooling
Exploratory testing
Python
Playwright
Statistics
Documentation

Tools

Playwright

Job description

Newton Research, a Boston-area AI-first QA startup, seeks a QA Engineer focused on evals for AI-native agents. You will own curated eval datasets, scoring rubrics, and regression tracking across model, prompt, skill, and agent releases to ensure safe, reliable agent behavior.

You will test changes before release, validate tool-call correctness, and cover end-to-end agent flows. The role emphasizes non-determinism handling, data-driven decisions, and collaboration with engineering on testability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI-Native QA Engineer for LLM/Agent Quality
AI-Native QA Engineer for LLM/Agent Quality

Newton Research • Massachusetts

On-site
USD 115,000 - 130,000
Salary equity
QA Engineer-AI Native Quality
QA Engineer-AI Native Quality

TestDevJobs • Boston (MA), Northern (KY)

Hybrid
USD 115,000 - 130,000
QA Engineer-AI Native Quality
QA Engineer-AI Native Quality

Newton Research • Massachusetts

On-site
USD 115,000 - 130,000
Salary equity
QA Engineer - Agentic Systems
QA Engineer - Agentic Systems

Meet Life Sciences • New York (NY)

On-site
USD 110,000 - 170,000
Remote AI Evaluation & Quality Assurance Engineer
Remote AI Evaluation & Quality Assurance Engineer

Equiliem • United States

Remote
USD 69,000 - 74,000
AI QA Engineer: Test Intelligent AI Features
AI QA Engineer: Test Intelligent AI Features

STATE STREET CORPORATION • Burlington (MA)

On-site
USD 80,000 - 140,000
401K with company match
Medical, dental, vision insurance
Paid time off
AI QA Validation Lead: Data Integrity for AI Initiatives
AI QA Validation Lead: Data Integrity for AI Initiatives

Teradyne • Reading (MA)

On-site
USD 126,000 - 202,000
Health benefits
Retirement plans
Tuition assistance
AI QA Engineer
AI QA Engineer

Cavendish Professionals • Town of Italy (NY)

On-site
USD 95,000 - 120,000
AI Quality Engineer: Automate Trust in Agent-Powered Systems
AI Quality Engineer: Automate Trust in Agent-Powered Systems

Dynasty Financial Partners • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
Remote AI-First QA Lead for Reliable Releases
Remote AI-First QA Lead for Reliable Releases

Jobgether SRL • United States

Remote
USD 140,000 - 180,000
Fully remote
Competitive salary
Vacation 28 days per year
+7