Job Description - QA Engineer - AI Platforms & Enterprise QA
Experience: 8-10 years
Role: Individual Contributor
Reports to: Product & Engineering Head
Role Summary
Ensure quality, reliability and production readiness of enterprise AI platforms, including multimodal/video-intelligence workflows, through functional, API, integration, workflow, modelevaluation and regression testing. Own test design, test-case quality, automation and AI-output validation across the product lifecycle.
Key Responsibilities
- Develop end-to-end QA strategy, test plans and release-quality criteria for enterprise AI platform capabilities.
- Design, document, execute and maintain functional, API, integration, workflow, regression and end-toend test cases.
- Translate product requirements, user stories and acceptance criteria into positive, negative, boundary, failure and edge-case test scenarios.
- Validate AI workflows, business rules, orchestration logic, exception handling and human-in-the-loop flows.
- Build and maintain automated regression suites and integrate them into CI/CD pipelines.
- Perform API validation across internal services, model endpoints and third-party integrations.
- Validate performance, scalability, reliability, retries, timeouts and graceful fallback behavior.
- Work closely with Product, Engineering and Architecture teams to identify quality risks early in the development cycle.
- Track, triage and drive defects to closure with clear reproduction steps, severity and impact assessment.
Test Design & Test Case Ownership
- Own the test-case repository, ensuring traceability from requirements and user stories to test scenarios, expected results and release sign-off.
- Define reusable test cases for APIs, workflows, role/tenant controls, model orchestration, model disagreement, confidence thresholds and fallback paths.
- Create representative test data covering normal, difficult and adversarial scenarios, including lowquality or ambiguous inputs.
- Maintain regression suites from approved test cases and ensure critical user journeys are automated wherever practical.
- Define clear pass/fail criteria for deterministic software behavior as well as probabilistic AI outputs.
AI Quality & Evaluation Responsibilities
- Design and maintain benchmark datasets and controlled golden evaluation sets for repeatable AI quality testing.
- Validate multimodal AI outputs across video, image, audio, ASR, OCR, metadata extraction, tagging, semantic search and retrieval use cases, as applicable.
- Compare outputs across multiple models and validate model-routing/orchestration decisions, including agreement, disagreement and fallback scenarios.
- Test prompts, structured outputs, confidence thresholds and confidence calibration; assess false positives, false negatives and hallucination risks.
- Establish and validate ground-truth reference data using authoritative metadata and/or human-reviewed annotations.
- Run regression evaluations whenever models, prompts, thresholds, workflows or orchestration policies change.
- Measure AI quality using appropriate metrics such as precision, recall, F1, WER, retrieval relevance, timestamp accuracy and task-specific acceptance criteria.
- Validate human-in-the-loop review, correction and adjudication workflows.
- Assess latency, throughput, reliability and inference-cost impact as part of production-readiness testing.
Required Technical Skills
- API testing: Postman, Swagger/OpenAPI or equivalent.
- Programming/scripting: Python preferred; Java acceptable.
- SQL and structured/unstructured data validation.
- Test automation frameworks and reusable test-fixture/test-data design.
- AI/ML evaluation fundamentals: ground truth, golden datasets, precision, recall, F1, false positives/negatives and confidence thresholds.
- LLM and multimodal AI validation: prompts, model outputs, hallucination/error analysis, semantic relevance and response consistency.
- Video intelligence testing: scene/event detection, OCR, ASR, logo/object/person detection, timestamps and semantic search.
- Multi-model evaluation: compare outputs from multiple models and validate routing, fallback and disagreement-handling logic.
- Regression evaluation: ability to create repeatable benchmark tests across model, prompt and orchestration changes.
- Performance testing: latency, throughput, concurrency and model/API response-time testing.
- AI observability: ability to inspect model calls, prompts, outputs, confidence scores, failures, retries and traces.
- CI/CD integration and quality gates for automated functional and AI regression suites.
Preferred Experience
- Enterprise AI platforms, SaaS products or workflow-automation platforms.
- Testing AI/ML or multimodal applications, including model APIs and probabilistic outputs.
- Video/media intelligence, content-processing, search/retrieval or metadata workflows is an advantage.
- Experience with benchmark datasets, golden sets, regression evaluation and production-quality monitoring is preferred.
Success Measures
- Low escaped production defects and strong release-quality predictability.
- High automation coverage for critical workflows and repeatable regression suites.
- Complete and maintainable test-case coverage with traceability to requirements and acceptance criteria.
- Reliable, measurable AI validation against agreed benchmark and golden-set criteria.
- Early detection of regressions across software, models, prompts and orchestration changes.
- Production readiness across functional quality, AI-output quality, performance, reliability and fallback behavior.