An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Wolters Kluwer - Financial Services Solutions seeks a quality engineering lead for enterprise-scale generative AI applications. You will own the AI quality framework, from RAG pipelines to LLM orchestration and AI-driven user experiences.
You will collaborate with developers to validate system architecture, observability, and production readiness while driving responsible AI practices. The role emphasizes building automated regression suites, harnesses for evaluation, and measurable gates for AI
Own and implement the quality engineering approach for enterprise-scale generative AI applications, including RAG pipelines, LLM orchestration, agentic workflows, APIs, ingestion pipelines, and AI-driven user experiences. Architect and implement evaluation harnesses that measure retrieval accuracy, citation correctness, answer relevance, groundedness, hallucination rate, safety, latency, cost, and end-to-end system behavior. Design automated regression suites that detect quality drift across prompts, embeddings, chunking strategies, model versions, retrieval configuration, orchestration logic, and application workflows. Partner with developers to validate AI system architecture, service integrations, API behavior, data ingestion, observability, resilience, security controls, and production readiness. Lead root‑cause analysis for AI quality failures using logs, traces, evaluation outputs, feedback signals, customer scenarios, and domain‑specific acceptance criteria. Define, implement, and maintain measurable quality gates for AI releases, and advise product and engineering leaders on risk, readiness, and trade‑offs across accuracy, reliability, latency, cost, and user trust. Develop test automation in Python and related frameworks for APIs, services, chat workflows, data pipelines, and AI evaluation workflows within CI/CD pipelines. Build and maintain observability practices and dashboards that track model behavior, retrieval performance, operational health, quality trends, and customer‑impacting failure modes. Contribute to responsible AI practices including bias and fairness checks, content safety validation, prompt injection testing, data privacy validation, and compliance‑oriented evidence collection. Collaborate with product managers, UX, subject matter experts, and stakeholders to translate customer workflows and business requirements into testable AI quality criteria. Mentor engineers and quality team members in AI evaluation methods, non‑deterministic testing strategies, automation design, and production‑quality engineering practices. Stay current with emerging AI quality, LLM evaluation, RAG assessment, and LLMOps practices through research, experimentation, and community engagement.