Fullstack AI Quality Engineer – Applied & Agentic AI Systems

Roche Holding AG

Hyderabad

On-site

INR 1,500,000 - 2,300,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Roche Holding AG à Hyderabad recherche un ingénieur QA IA expérimenté pour tester des applications IA Génératives et agentiques. Vous couvrirez les tests manuels et automatisés sur Python et TypeScript, évaluant les performances, l’éthique et la sécurité des systèmes IA, en accord avec les exigences réglementaires et les pipelines CI/CD.

Le candidat idéal possède une solide expérience en tests IA, une connaissance des cadres de test (pytest, allure, Playwright, Cypress) et une sensibilité aux

Qualifications

  • 7+ années d'expérience en tests logiciels, au moins 1 an en genAI/applications agentiques.
  • Expérience en environnements régulés avec CSV/ITSM et compréhension de GxP.
  • Expérience dans des équipes Agile et intégration continue.
  • Expérience en tests IA et applications GenAI/agentiques.

Responsibilities

  • Concevoir et mettre en œuvre des stratégies de test pour les applications LLM et systèmes agentiques, évaluer les réponses et les pipelines RAG.
  • Concevoir et déployer des solutions de test automatisé pour l'évaluation continue des apps LLM sur backend Python et frontend TS.
  • Utiliser les cadres de test Roche et produire documentation et guidelines pour l'IA.
  • Concevoir des tests de cas limites, biais et équité dans les sorties IA.
  • Collaborer sur les tests de sécurité et de performance des applications GenAI.

Skills

Full-Stack Testing
Python Testing
TypeScript Testing
End-to-End Automation
CI/CD Pipelines
AI Quality Evaluation
Agentic Testing
Evaluation Datasets
Observability
AI Knowledge
Python Scripting

Education

BSc/BE/MSc in CS/SE

Tools

pytest
allure
Playwright
Jest
Cypress

Job description

Chez Roche, vous pouvez être vous-même et être apprécié pour les qualités uniques que vous apportez. Notre culture encourage l'expression personnelle, le dialogue ouvert et les connexions authentiques, où vous êtes valorisé, accepté et respecté pour ce que vous êtes, vous permettant de prospérer tant personnellement que professionnellement. Voici comment nous visons à prévenir, arrêter et guérir les maladies et à garantir à chacun l'accès aux soins de santé aujourd'hui et pour les générations à venir. Rejoignez Roche, où chaque voix compte.

La position

Job description

We are seeking a highly skilled AI Quality Engineer specializing in testing Generative (and non-Generative) AI applications and agentic systems to join our team. The ideal candidate will have expertise in both traditional software testing and specialized testing methodologies for Large Language Models (LLMs)-based and AI-based systems.

As a fullstack Quality Engineer within a lean, agentic software team, you will cover the entire spectrum of quality assurance. You will span manual exploratory testing, traditional automated UI/API/MCP testing (across our Python and TypeScript stack), classical machine learning validation, and cutting-edge generative and agentic AI evaluation. This role requires a deep understanding of LLM, agentic and ML systems behaviors, and testing frameworks to ensure the reliability, fairness, and effectiveness of our AI-powered solutions.

Description of the area

At Roche Digital Technology, we are advancing the boundaries of Applied AI. The Applied AI Engineering Team focuses on architecting, building, and operating high-value AI solutions and services to solve complex business challenges in healthcare.

In the 2026 tech landscape, we operate in small, highly autonomous agile teams (e.g., 9 members) powered by advanced coding agents (like Claude Code) to develop and ship solutions faster than ever before. In this highly regulated environment, quality cannot be an afterthought. You will be the foundational pillar ensuring our rapidly developed agentic workflows and AI, GenAI, and agentic applications are safe, compliant, and robust before they reach the clinical or enterprise user.

Job Responsibilities

  • Generative & Agentic AI Testing Strategy: Design and implement comprehensive testing strategies specifically tailored for LLM-based applications and agentic systems, including evaluation of model responses, RAG pipelines accuracy, and overall system reliability. You will expand this to evaluate autonomous multi-agent behaviors and tool-use accuracy.

  • Full-Stack Test Automation: Design and implement automated testing solutions for continuous evaluation of LLM applications, including integration tests, performance tests, and specialized AI behavior tests. This spans both our Python backend services and modern TypeScript/React/Angular frontends.

  • Quality Assurance Framework & Compliance: Utilize Roche's testing frameworks that address both traditional software quality aspects and AI-specific concerns such as output consistency, contextual accuracy, and ethical compliance; co-create and maintain such frameworks when required. Create and maintain comprehensive test documentation testing strategy, test cases, and testing guidelines specific to AI applications and compliant with Roche practices.

  • Edge Case, Bias, and Fairness Testing: Identify and develop test scenarios for edge cases in LLM behavior, including handling of ambiguous inputs, potential biases, and unexpected response patterns. Design and execute tests to identify potential biases in model outputs and ensure fair treatment across different user groups and use cases.

  • Security & Performance Testing: Collaborate within development teams to test for potential vulnerabilities specific to LLM applications, including prompt injection, data leakage, and other AI-specific security concerns. Collaborate with other engineers to conduct thorough performance testing of GenAI applications, including response time analysis, load testing, and resource utilization monitoring.

Qualifications

  • Hold B.Sc., B.Eng., M.Sc., M.Eng. in Computer Science, Software Engineering, or equivalent degree with strong background in testing methodologies.

  • Be passionate about quality assurance and testing, particularly in the context of AI and GenAI applications and agents.

  • Have strong analytical and problem-solving skills with attention to detail.

  • Utilize critical thinking and problem-solving when using AI prompts (in e.g. NLP, data distributions and edge cases).

  • Be proactive in identifying potential issues and suggesting improvements.

Education / Experience

  • Experience: 7+ years of experience in software testing, with at least 1 year focused on genAI and/or agentic applications.

  • Regulated Environments: Be ready to work in a regulated environment with qualified infrastructure and validated applications, understand the difference between GxP and non-GxP products and the importance of CSV and ITSM processes in the delivery.

  • Agile Integration: Be experienced in working with Agile development teams.

Skills

Must-have:

  • Comprehensive Full-Stack Testing Mastery

  • Core Languages & Testing Frameworks: Advanced command of Python and standard testing frameworks (e.g., pytest, allure), alongside practical familiarity with specialized AI evaluation tools. High competence in TypeScript-driven test frameworks (such as Playwright, Jest, and Cypress) for thoroughly verifying dynamic generative user interfaces.

  • End-to-End Automation & CI/CD: Proven expertise in architecting and maintaining end-to-end automated testing suites that bridge traditional UI/API validation with complex, asynchronous AI/LLM evaluation pipelines.

  • Pipeline Integration: Hands-on experience integrating these comprehensive suites within GitHub Actions or modern CI/CD pipelines to ensure continuous, reliable quality checks throughout the entire development lifecycle.

  • Advanced AI Quality & Evaluation Mastery

  • AI Evaluation Frameworks: Expertise in implementing evaluation harnesses (e.g., RAGAS, DeepEval, Promptfoo) to quantify LLM performance, hallucination rates, and answer relevancy.

  • Agentic Workflow Validation: Ability to test autonomous agentic loops, including tool-use accuracy, step-by-step reasoning verification, and goal-directed performance metrics.

  • Evaluation Datasets: Experience in curating, managing, and versioning 'Golden Datasets' for benchmarking model responses against ground truth in both development and production environments.

  • Observability & Feedback Loops: Leveraging tracing (e.g., LangSmith, Arize Phoenix) to monitor production inputs/outputs, identify drift, and capture user feedback for continuous improvement.

  • Basic broad AI knowledge: Understanding of LLM architectures, RAG systems, agents, and common failure modes in AI applications.

  • Python Expertise: Advanced proficiency in writing robust, maintainable Python code for test automation (e.g., pytest). Experience using Python to orchestrate complex AI evaluation workflows, process large datasets for benchmarking, and automate interactions with LLM/AI services. Ability to integrate Python scripts into CI/CD pipelines for automated quality checks and performance observability in agentic systems.

Should-have:

  • Cloud: Practical experience with testing applications on cloud platforms (AWS) and working with cloud-based AI services.

  • Observability & ALM: Experience with monitoring and observability tools, log analysis, and performance metrics tracking (Grafana). Experienced in Roche mandatory ALM tools (Jira, GitHub).

  • Agentic SDLC & Engineering Excellence: Leverage AI coding assistants and autonomous agents (e.g., Claude Code, Ona) daily to accelerate full-stack development and testing cycles. Conduct rigorous code reviews for both human-written and AI-generated code.

  • Security: Understanding of security testing methodologies, particularly in the context of AI applications.

Could-have:

  • Regulatory Compliance: Proven understanding of AI quality assurance and testing standards within highly regulated industries.

Additional Qualifications

  • Stay current with the latest developments in AI and GenAI testing methodologies and best practices.

  • Have excellent communication skills and the ability to clearly document findings and recommendations.

  • Be able to communicate in English at the level of C1+.

  • Located in Hyderabad, India, with working hours structured to capture the 'golden hours' of overlap with Central European Time (typically running through the IST evening).

#Hyd2026

Qui nous sommes

Un avenir plus sain nous pousse à innover. Ensemble, plus de 100 000 employés à travers le monde sont dédiés à faire progresser la science et à garantir à chacun l'accès aux soins de santé aujourd'hui et pour les générations à venir. Nos efforts aboutissent à plus de 26 millions de personnes traitées avec nos médicaments et plus de 30 milliards de tests réalisés avec nos produits de Diagnostique. Nous nous encourageons mutuellement à explorer de nouvelles possibilités, à favoriser la créativité et à conserver nos grandes ambitions, afin de fournir des solutions de santé qui changent des vies et ont un impact mondial.

Construisons ensemble un avenir plus sain.

Roche est un employeur offrant l'équité en matière d'emploi.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Fullstack Data Engineer - Applied & Agentic AI Systems
Fullstack Data Engineer - Applied & Agentic AI Systems

Roche Holding AG • Hyderabad

On-site
INR 3,500,000 - 6,000,000
Aucune mention spécifique de perks
AI Solutiuons Architect
AI Solutiuons Architect

Roche Holding AG • Hyderabad

On-site
INR 4,000,000 - 7,500,000
Environnement innovant
Opportunités de croissance
Culture collaborative
Head of Forward-Deployed Data Engineering, India Center
Head of Forward-Deployed Data Engineering, India Center

Roche Holding AG • Hyderabad

On-site
INR 4,000,000 - 8,000,000
Software Developer Regulatory Affairs and China Clinical Development - Pharma R&D
Software Developer Regulatory Affairs and China Clinical Development - Pharma R&D

Roche Holding AG • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Roche benefits
Fullstack UX/UI Researchers/Designer - Applied & Agentic AI Systems
Fullstack UX/UI Researchers/Designer - Applied & Agentic AI Systems

Roche Holding AG • Hyderabad

On-site
INR 400,000 - 700,000
Data Scientist f Submission Data and Content Generation & Reuse (AIDCG) - Pharma R&D
Data Scientist f Submission Data and Content Generation & Reuse (AIDCG) - Pharma R&D

Roche Holding AG • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Patient Strategy AWS Tech Lead
Patient Strategy AWS Tech Lead

Roche Holding AG • Hyderabad

On-site
INR 4,500,000 - 7,500,000
Fullstack AI Quality Engineer – Applied & Agentic AI Systems
Fullstack AI Quality Engineer – Applied & Agentic AI Systems

F. Hoffmann-La Roche AG • Hyderabad

On-site
INR 250,000 - 450,000
Patient Strategy - Integration Tech Lead
Patient Strategy - Integration Tech Lead

Roche Holding AG • Hyderabad

On-site
INR 2,500,000 - 4,200,000
Fullstack AI Quality Engineer – Applied & Agentic AI Systems
Fullstack AI Quality Engineer – Applied & Agentic AI Systems

Roche • Hyderabad

On-site
INR 1,500,000 - 2,600,000