QA Engineer-AI Native Quality

TestDevJobs

Boston, Northern (MA, KY)

Hybrid

USD 115,000 - 130,000

Full time

19 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Newton Research, a Boston-area AI-first QA startup, seeks a QA Engineer focused on evals for AI-native agents. You will own curated eval datasets, scoring rubrics, and regression tracking across model, prompt, skill, and agent releases to ensure safe, reliable agent behavior.

You will test changes before release, validate tool-call correctness, and cover end-to-end agent flows. The role emphasizes non-determinism handling, data-driven decisions, and collaboration with engineering on testability.

Qualifications

  • 4+ years in QA or test engineering on a complex web product.
  • Hands-on experience building evals for LLM or agent products (datasets, rubrics, LLM-as-judge, regression tracking), or a clear track record of getting there fast
  • Working knowledge of how agents work: prompting, skills, tool calling, context, RAG and orchestration, well enough to tell where a failure originates
  • An automation-first instinct: you can show what you removed from a manual process
  • Python (or similar) for eval tooling and data checks; Playwright or similar for end-to-end tests
  • Statistical literacy: pass rates, variance and sample size when outputs are not deterministic
  • Daily user of AI coding and testing assistants, with the judgment to verify their output rather than trust it
  • Strong exploratory instincts and excellent written communication
  • Nice to have: adtech, martech or marketing analytics domain; SSO / SAML / OAuth flows; data-connector testing; eval or observability tooling

Responsibilities

  • Own the eval suite for Newton's agents: curated datasets of inputs and expected behaviors, rubric and LLM-as-judge scoring, regression tracking across model, prompt, skill and agent releases; validate judges against human-labeled examples
  • Test skill and prompt changes before they ship: every edit to a skill, system prompt or tool definition runs against the relevant evals in CI, with a before/after comparison a reviewer can read in one glance; gate releases on the results
  • Cover the agent flows end to end: tool-call correctness, task completion, multi-turn coherence, blueprint creation from conversations, code generation, scheduled tasks; flag output that is wrong, empty or silently degraded
  • Handle non-determinism with rigor: repeated runs, pass-rate thresholds and simple statistics, so a flaky agent is a measured finding, not an anecdote
  • Turn production and customer signal into evals: mine logs, error tracking and customer reports so every escaped bad-behavior case becomes a permanent eval, ideally drafted by an agent and reviewed by you
  • Probe AI-specific risk: prompt injection, data leakage across users, projects and permissions, hallucinated or ungrounded numbers in analytics output, cost and latency regressions
  • Automate before you repeat: any check you do by hand twice becomes a Playwright test, an eval or an agent workflow; drive manual regression time down every sprint
  • Design agentic QA workflows in CI: agents that run suites, triage failures, draft defects and re-verify fixes, with guardrails, cost limits and escalation rules you define
  • Keep human-judgment work sharp: exploratory testing across roles, feature flags and environments, SSO and connector authorization, and the cross-feature bugs only a person thinks to look for
  • Write bug reports that are machine- and human-readable, and partner with engineering on testability (observability, seedable data, stable interfaces)

Skills

QA testing
Eval tooling
Exploratory testing
Python
Playwright
Statistics
Documentation

Tools

Playwright

Job description

QA Engineer, AI-Native Quality
Newton Research · Research & Development · Boston / Needham, MA

Company Description

Newton Research is a fast-growing software start-up founded by repeat entrepreneurs and well-funded by blue chip venture capital firms. We are building the next generation of the closed loop media lifecycle, developing AI agents that leverage the latest in LLMs and generative AI with specialized knowledge. Our products generate actionable business insights for our customers and partners, assisting in each step of the media planning, buying and measurement lifecycle.

About the Role

Newton ships on a sprint cadence through a develop, stage and customer-environment pipeline, and the product surface is wide: conversations, blueprints, connectors, scheduled tasks, permissions and sharing, SSO, and AI agents whose behavior is not fully deterministic. A missed regression lands in front of a media planner or a customer's security review.
We run everything through an AI-first lens, because it is the only way quality scales. If a quality task is repeatable, an agent does it and you supervise; if it takes judgment, that is where you spend your time.
The gap this hire fills: evals and skill-change testing. Code changes already have CI and review, including PRs written by Claude. What has no safety net is behavior change: an edit to a skill, a prompt, a tool definition or a model version can silently change what our agents do, and nothing tests that today. You own that layer: curated eval sets, scoring and regression tracking, so any change to how an agent behaves is measured before it ships.

You are also a release-readiness partner (are we good?), alongside our existing QA lead, and that judgment stays human. But you are not hired to test Claude-driven PRs line by line.

What You Will Do
  • Own the eval suite for Newton's agents: curated datasets of inputs and expected behaviors, rubric and LLM-as-judge scoring, regression tracking across model, prompt, skill and agent releases; validate judges against human-labeled examples
  • Test skill and prompt changes before they ship: every edit to a skill, system prompt or tool definition runs against the relevant evals in CI, with a before/after comparison a reviewer can read in one glance; gate releases on the results
  • Cover the agent flows end to end: tool-call correctness, task completion, multi-turn coherence, blueprint creation from conversations, code generation, scheduled tasks; flag output that is wrong, empty or silently degraded
  • Handle non-determinism with rigor: repeated runs, pass-rate thresholds and simple statistics, so a flaky agent is a measured finding, not an anecdote
  • Turn production and customer signal into evals: mine logs, error tracking and customer reports so every escaped bad-behavior case becomes a permanent eval, ideally drafted by an agent and reviewed by you
  • Probe AI-specific risk: prompt injection, data leakage across users, projects and permissions, hallucinated or ungrounded numbers in analytics output, cost and latency regressions
  • Automate before you repeat: any check you do by hand twice becomes a Playwright test, an eval or an agent workflow; drive manual regression time down every sprint
  • Design agentic QA workflows in CI: agents that run suites, triage failures, draft defects and re-verify fixes, with guardrails, cost limits and escalation rules you define
  • Keep human-judgment work sharp: exploratory testing across roles, feature flags and environments, SSO and connector authorization, and the cross-feature bugs only a person thinks to look for
  • Write bug reports that are machine- and human-readable , and partner with engineering on testability (observability, seedable data, stable interfaces)
What Makes You a Great Fit
  • 4+ years in QA or test engineering on a complex web product, ideally B2B SaaS shipping frequently
  • Hands-on experience building evals for LLM or agent products (datasets, rubrics, LLM-as-judge, regression tracking), or a clear track record of getting there fast
  • Working knowledge of how agents work: prompting, skills, tool calling, context, RAG and orchestration, well enough to tell where a failure originates
  • An automation-first instinct: you can show what you removed from a manual process
  • Python (or similar) for eval tooling and data checks; Playwright or similar for end-to-end tests
  • Statistical literacy: pass rates, variance and sample size when outputs are not deterministic
  • Daily user of AI coding and testing assistants, with the judgment to verify their output rather than trust it
  • Strong exploratory instincts and excellent written communication
  • Nice to have: adtech, martech or marketing analytics domain; SSO / SAML / OAuth flows; data-connector testing; eval or observability tooling
How Success Is Measured
  • Eval suite covering Newton's core agent flows, run on every model, prompt, skill or agent change, with judge accuracy checked against human labels
  • Share of skill and prompt changes that ship with a before/after eval result (target: all of them)
  • Escaped bad-behavior cases converted into permanent evals
  • Share of regression coverage run automatically or by agents, rising every sprint
  • Release calls that hold up: few surprises after release

Salary range: $115,000-130,000 + Equity

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

QA Engineer-AI Native Quality
QA Engineer-AI Native Quality

Newton Research • Massachusetts

On-site
USD 115,000 - 130,000
Salary equity
QA Engineer: AI-Native Quality, Equity & Impact
QA Engineer: AI-Native Quality, Equity & Impact

TestDevJobs • Boston (MA), Northern (KY)

Hybrid
USD 115,000 - 130,000
Member of Technical Staff (QA Engineer - Agentic Systems)
Member of Technical Staff (QA Engineer - Agentic Systems)

Solstice • New York (NY)

On-site
USD 160,000 - 300,000
Health, dental, and vision insurance
Ground-floor equity opportunity
401(k) w/ match
+3
QA AI Automation Engineer
QA AI Automation Engineer

Dynasty Financial Partners • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
Project Lead, AI Model Training
Project Lead, AI Model Training

NewtonX • New York (NY)

Hybrid
USD 150,000 - 240,000
Competitive compensation
Stock options
Health benefits
+2
QA AI Automation Engineer
QA AI Automation Engineer

Dynasty Financial Partners • Town of Florida (NY), Northern (KY)

Hybrid
USD 120,000 - 150,000
AI-Native QA Engineer for LLM/Agent Quality
AI-Native QA Engineer for LLM/Agent Quality

Newton Research • Massachusetts

On-site
USD 115,000 - 130,000
Salary equity
QA AI Automation Engineer
QA AI Automation Engineer

Dynasty Financial Partners, LLC • Saint Petersburg (FL)

On-site
USD 120,000 - 150,000
QA Engineer
QA Engineer

Neuroscale AI • United States

On-site
USD 90,000 - 115,000
Medical, dental, and vision benefits
Professional growth opportunities
Remote-first / hybrid flexibility
QA Engineer - Agentic Systems
QA Engineer - Agentic Systems

Meet Life Sciences • New York (NY)

On-site
USD 110,000 - 170,000