Agentic AI Quality Engineer

Accenture the Netherlands

Amsterdam

Hybrid

EUR 90,000 - 120,000

Full time

17 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

4x9 Workweek Option
Hybrid Working
FlexDay
Competitive Base Salary
MyBenefits Budget
Annual Bonus
Net Allowance
Employee Share Purchase Plan
Collective Health Insurance Scheme
Mental Health Support
Well-being Hub
26 Vacation Days
Special Leave Types
Culture Days
NS Business Card
Mobility Budget
Performance Achievement Program
Learning & Development Platform
Guidance from People Leads
Do Good Volunteer Program
I&D Networks
Refugee Talent Track
Accentrics & Friday Drinks
Annual Accenture NL Party
Flat Organizational Hierarchy

Job summary

Accenture the Netherlands is seeking a Senior ML QA Engineer to lead quality and performance evaluation across conversational AI and multi-step workflows. You will work with quality engineers, software engineers, AI specialists, and business consultants to embed testing in CI/CD, monitor telemetry, and drive safe, compliant decisions for client projects.

You will contribute to governance, observability, and evaluation strategy while mentoring teams and sharing reusable assets across engagements.

Qualifications

  • 3+ years in machine learning QA, AI evaluation engineering, or MLOps engineering.
  • Bachelor’s or Master’s degree in Computer Science, Data Science, or a related quantitative field.
  • Dutch language proficiency (B2+) for collaboration with Dutch-speaking clients and stakeholders; English at C2 level.
  • Strong knowledge of prompt engineering, evaluation tools, and telemetry.

Responsibilities

  • Automate evaluation of multi-turn conversations and long-horizon tasks using metrics for reasoning accuracy, evidence grounding, plan execution, latency, token usage, and recovery.
  • Apply schema validation, tracing, and calibrated judges to detect inaccurate tool calls, loops, and misuse.
  • Run red-team, prompt-injection, jailbreak, privacy, and PII-leak tests across agent workflows.
  • Embed test harnesses in CI/CD and monitor production telemetry for regressions, behavioral drift, and recovery.
  • Translate business risks into measurable quality criteria, practical controls, and release recommendations.
  • Communicate findings and residual risks for go/no-go decisions; coach teams and share reusable assets and standards.

Skills

Agentic AI
Python
Prompt engineering
Vector databases
Execution telemetry
Security mindset
Leadership

Education

Bachelor’s or Master’s in Computer Science, Data Science, or related quantitative field

Tools

MLflow
Galileo
PyRIT

Job description

What you’ll do:
Quality & Performance Evaluation
  • Automate evaluation of multi-turn conversations and long-horizon tasks using metrics for reasoning accuracy, evidence grounding, plan execution, latency, token usage, and recovery.
  • Apply schema validation, tracing, and calibrated LLM-as-a-Judge frameworks to detect inaccurate tool calls, loops, and misuse.
Safety & Compliance Guardrails
  • Implement filters and golden datasets to test harmful content, policy violations, hallucinations, bias, and brand drift.
  • Use human-in-the-loop review for high-risk edge cases.
Adversarial Security & Red Teaming
  • Run red-team, prompt-injection, jailbreak, privacy, and PII-leak tests across agent workflows.
  • Analyze failures in model logic and API integrations to identify root causes and controls.
MLOps & Production Monitoring
  • Embed test harnesses in CI/CD and monitor production telemetry for regressions, behavioral drift, and recovery.
  • Maintain evaluation data lineage, versioning, curation, and golden responses.
Consulting & Collaboration
  • Translate business risks into measurable quality criteria, practical controls, and release recommendations.
  • Align stakeholders on user journeys, edge cases, thresholds, mitigations, testability, observability, governance, and evaluation strategy.
  • Communicate findings and residual risks for go/no-go decisions; coach teams and share reusable assets and standards.
What Success Looks Like (KPIs)
  • Safety: Zero critical production incidents.
  • Automation: High coverage of non-deterministic agent paths.
  • Speed: Rapid regression detection and remediation after model updates.
Industry Context & Methodology Alignment
  • Treat non-determinism as a core characteristic. As the Data Science Collective notes, robust testing verifies correct reasoning, grounded evidence, scoped actions, and safe, compliant decisions.
  • Drive continuous optimization by embedding evaluation in CI/CD, using tunable judges, aligned scorers, automated dataset updates, and dashboards, as outlined by Databricks.
How We Reinvent

At Accenture, our teams operate within one of our core service lines— Digital Core, Industry & Enterprise, Supply Chain & Engineering, Cybersecurity, Song, Finance, Talent, Client Groups, Global Operations, Reinventions Services —each designed to drive reinvention at scale. This role falls under Digital Core, where we help clients modernize their technology foundations and accelerate business transformation through cloud, data, software engineering, and AI.

You will join multidisciplinary teams consisting of quality engineers, software engineers, architects, AI specialists, platform engineers, and business consultants. Together, we help organizations reinvent how software is designed, built, tested, deployed, and operated through the power of AI.

We go to market by industry, leveraging our global delivery network and strategic partnerships with ecosystem leaders like Microsoft, AWS, and Google. Our culture of shared success and continuous learning ensures that every team member contributes to—and benefits from—our reinvention journey.

What we’re looking for:
  • Education: Bachelor’s or Master’s in Computer Science, Data Science, or a related quantitative field.
  • Experience: 3+ years in machine learning QA, AI evaluation engineering, or MLOps engineering.
  • Language(s): Dutch – B2+ for collaboration with Dutch-speaking clients and stakeholders.
  • Language(s): English – C2.
  • Skills: Agentic AI Expertise: Experience evaluating autonomous agents, multi-step workflows, or Retrieval-Augmented Generation (RAG) applications.
  • Skills: Technical Stack: Strong Python proficiency and experience with AI evaluation tools such as MLflow, Galileo, and PyRIT.
  • Skills: Framework Mastery: Deep knowledge of prompt engineering, vector databases, tool schema design, and execution telemetry.
  • Skills: Security Mindset: Hands-on experience in adversarial testing, threat modeling, or model safety guardrails.
  • Skills: Leadership: Strong communication, technical leadership, a growth mindset, and commitment to continuous innovation.
Bonus points for
  • TMAP or ISTQB Certification
  • Contributions to open-source projects or engineering communities
What we offer:
  • Work-Life Flexibility: 4x9 Workweek Option, Hybrid Working, FlexDay
  • Financial Well-being: Competitive Base Salary, MyBenefits Budget, Annual Bonus, Net Allowance, Employee Share Purchase Plan
  • Health & Wellness: Collective Health Insurance Scheme, Mental Health Support, Well-being Hub
  • Time Off & Leave: 26 Vacation Days, Special Leave Types (parental leave, partner leave, bereavement leave, and gender affirmation leave), Culture Days
  • Mobility & Commuting: NS Business Card, Mobility Budget (option to choose between electric bike, car lease, or public transport reimbursement)
  • Career Development: Performance Achievement Program, Learning & Development Platform, Guidance from People Leads
  • Inclusion, Diversity & Social Impact: Do Good Volunteer Program, I&D Networks, Refugee Talent Track
  • Culture & Community: Accentrics & Friday Drinks, Annual Accenture NL Party, Flat Organizational Hierarchy

Ready to lead and innovate?

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Agentic AI Quality Engineer
Agentic AI Quality Engineer

Accenture • Amsterdam

On-site
EUR 90,000 - 140,000
Agentic AI Quality Engineer
Agentic AI Quality Engineer

parsionate • Amsterdam

On-site
EUR 90,000 - 120,000
AI Security (Senior) Consultant
AI Security (Senior) Consultant

Accenture • Amsterdam

On-site
EUR 90,000 - 120,000
Work-Life Flexibility
Financial Well-being
Health & Wellness
+5
AI Security (Senior) Manager – Architecture/Deliver
AI Security (Senior) Manager – Architecture/Deliver

Accenture the Netherlands • Amsterdam

Hybrid
EUR 120,000 - 180,000
4x9 workweek
Hybrid working
Annual bonus
+5
AI Security (Senior) Consultant
AI Security (Senior) Consultant

Accenture the Netherlands • Amsterdam

Hybrid
EUR 90,000 - 150,000
4x9 Workweek
Hybrid Working
FlexDay
+23
AI Security (Senior) Manager – Architecture/Deliver
AI Security (Senior) Manager – Architecture/Deliver

Accenture • Amsterdam

On-site
EUR 120,000 - 180,000
Work‑Life Flexibility: 4x9 Workweek
Hybrid Working
Annual Bonus
+4
Senior AI-Native Java Engineer
Senior AI-Native Java Engineer

Accenture • Amsterdam

Hybrid
EUR 90,000 - 140,000
4x9 Workweek
Hybrid Working
Annual Bonus
+3
Senior AI-Native Java Engineer
Senior AI-Native Java Engineer

Accenture the Netherlands • Amsterdam

Hybrid
EUR 120,000 - 150,000
4x9 Workweek Option
Hybrid Working
Annual Bonus
+6
Data Science & Machine Learning Engineering Consultant – Strategy & Consulting
Data Science & Machine Learning Engineering Consultant – Strategy & Consulting

Accenture the Netherlands • Amsterdam

On-site
EUR 70,000 - 100,000
Flexible working hours
Parental leave
Pension scheme
Junior Consultant AI Strategy – Strategy & Consulting
Junior Consultant AI Strategy – Strategy & Consulting

Accenture • Amsterdam

On-site
EUR 60,000 - 85,000
Unlimited learning
Flexible working hours
Paid transport
+4