AI QA Engineer for LLM & GenAI Testing & Automation

Compunnel, Inc.

New York, Northern (NY, KY)

Hybrid

USD 140,000 - 190,000

Full time

12 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Compunnel, Inc. seeks an Evaluation Engineer - AI Models in New York to evaluate and validate AI/ML and Generative AI models across business and technical use cases.

This role combines extensive Quality Engineering and software testing experience with hands-on AI model evaluation, production-grade LLM testing, test automation, and data-driven analysis. Strong Selenium and Playwright expertise is required, along with excellent communication skills and the ability to collaborate with engineering,

Qualifications

  • 10+ years of experience in Quality Assurance, Quality Engineering, Software Testing, or related disciplines.
  • 2+ years of hands-on experience in AI model evaluation, Generative AI testing, or ML validation.
  • Strong hands-on experience with Selenium and Playwright.
  • Strong understanding of AI/ML concepts, LLM behavior, prompt evaluation, and model testing methodologies.
  • Experience testing production-grade LLMs and AI-powered applications.
  • Experience with API testing, test automation frameworks, and data validation techniques.
  • Understanding of evaluation metrics such as precision, recall, accuracy, grounding, relevance, and hallucination detection.
  • Experience creating automated and manual test strategies and frameworks.
  • Strong analytical and problem-solving skills.
  • Excellent verbal and written communication skills.
  • Ability to work independently and collaborate effectively with cross-functional teams.

Responsibilities

  • Evaluate and validate AI models across business and technical use cases.
  • Design evaluation strategies, validate model accuracy and performance, detect hallucinations and other quality issues.
  • Develop automated and manual testing frameworks and data-driven analysis.

Skills

Quality Assurance
Quality Engineering
AI model evaluation
Test automation
Data validation
Analytical thinking
Communication skills
Cross-functional collaboration

Tools

Selenium
Playwright

Job description

Compunnel, Inc. seeks an Evaluation Engineer - AI Models in New York to evaluate and validate AI/ML and Generative AI models across business and technical use cases.

This role combines extensive Quality Engineering and software testing experience with hands-on AI model evaluation, production-grade LLM testing, test automation, and data-driven analysis. Strong Selenium and Playwright expertise is required, along with excellent communication skills and the ability to collaborate with engineering,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Automation QA w/ AI testing
Automation QA w/ AI testing

Compunnel, Inc. • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 190,000
Agentic AI QA Engineer: Python Automation & Evaluation
Agentic AI QA Engineer: Python Automation & Evaluation

Compunnel, Inc. • Atlanta (GA), Northern (KY)

Hybrid
USD 110,000 - 160,000
Senior AI Engineer — LLM Evaluation & Production Systems
Senior AI Engineer — LLM Evaluation & Production Systems

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Senior AI Engineer - Production LLM EvalOps
Senior AI Engineer - Production LLM EvalOps

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarking & Diagnostics

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
AI Quality Scientist - LLM Evaluation & Metrics (Hybrid)
AI Quality Scientist - LLM Evaluation & Metrics (Hybrid)

Programmers.io • Louisville (KY)

Hybrid
USD 90,000 - 130,000
QA Engineer - Agentic Systems
QA Engineer - Agentic Systems

Meet Life Sciences • New York (NY)

On-site
USD 110,000 - 170,000
Senior AI Engineer: LLM Evaluation & Production
Senior AI Engineer: LLM Evaluation & Production

LawPro.ai • Town of Florida (NY)

On-site
USD 140,000 - 210,000
QA / Automation Engineer Agentic AI
QA / Automation Engineer Agentic AI

Compunnel, Inc. • Atlanta (GA), Northern (KY)

Hybrid
USD 110,000 - 160,000