100% Remote – QA Automation OR Data Scientist with AI Exp.

SDH Systems

United States

Remote

USD 120,000 - 155,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

SDH Systems is seeking a Senior QA Automation Engineer / Data Scientist (Generative AI) for 100% remote work. The role focuses on evaluating AI-generated outputs, defining quality metrics, and testing non-deterministic systems.

Candidates from QA automation, data science, or ML backgrounds are welcome. The ideal candidate will research innovative evaluation methods, build automated testing pipelines, and collaborate across product, engineering, and data science teams to improve AI quality over

Qualifications

  • Background as a Data Scientist or QA Automation Engineer with related degree or experience.
  • Expertise and/or hands-on experience with LLM evaluation.
  • Proficient in Python or similar language used for testing/data analysis.
  • Strong analytical, research, and problem-solving abilities.

Responsibilities

  • Perform LLM evaluation and quality assurance for Generative AI applications.
  • Design and execute testing strategies for AI/LLM-powered systems.
  • Define evaluation criteria, quality standards, and success metrics for AI outputs.
  • Research innovative approaches for evaluating non-deterministic AI systems.
  • Create test datasets, benchmarks, and repeatable evaluation methodologies.
  • Identify quality issues such as hallucinations, inconsistencies, and inaccuracies.
  • Develop processes to measure AI quality over time.
  • Build and maintain automated testing/evaluation solutions.
  • Develop dashboards and reports communicating AI quality trends.
  • Collaborate with product owners, engineers, and data scientists.

Skills

Data Scientist
QA Automation Engineer
LLM evaluation
Python
Analytical thinking

Education

Degree in related field

Tools

DeepEval
OpenAI
Anthropic
Google Gemini

Job description

Job Title: Senior QA Automation Engineer / Data Scientist (Generative AI)

Location: 100% Remote

Interview Process: Video (2 Rounds)

Job Description:
We arelooking for a Data Scientist or QA Automation Engineer with expertise and/or experience in LLM (Large Language Model) evaluation to support quality assurance initiatives for Generative AI solutions. This role focuses on evaluating AI-generated outputs, defining quality metrics, and developing practical approaches for testing non-deterministic systems. Candidates from a variety of backgrounds – software testing, automation, data science, or machine learning – are encouraged to apply. The most important qualification for this role is the ability to research and develop innovative methods for evaluating LLM-based systems, where traditional pass/fail testing approaches may not be sufficient. Experience with DeepEval or similar evaluation frameworks is a plus, but not required.

Roles & Responsibilities
  • LLM Evaluation & Quality Assurance
  • Design and execute testing strategies for Generative AI and LLM-powered applications.
  • Define quality standards, evaluation criteria, and success metrics for AI-generated outputs.
  • Research and develop innovative approaches for evaluating non-deterministic AI systems.
  • Create test datasets, benchmark scenarios, and repeatable evaluation methodologies.
  • Identify and analyze quality issues including hallucinations, inconsistencies, inaccuracies, and reliability concerns.
  • Develop processes to measure and track AI quality over time.
Automation & Analysis
  • Build and maintain automated testing and evaluation solutions.
  • Analyze evaluation results and provide recommendations for quality improvement.
  • Support integration of AI quality validation into existing development and testing processes.
  • Develop reports, dashboards, and metrics that communicate AI quality and performance trends.
Collaboration & Innovation
  • Collaborate with product owners, engineers, data scientists, and business stakeholders to define quality expectations.
  • Support continuous improvement efforts through experimentation and data-driven analysis.
  • Stay current with emerging AI evaluation techniques, tools, and industry best practices.
  • Help establish repeatable quality assurance practices for AI-enabled solutions.
Required Qualifications
Must-Have Skills
  • Background as a Data Scientist, QA Automation Engineer, or similar technical role, with a degree in a related field or equivalent practical experience.
  • Expertise and/or hands-on experience with LLM (Large Language Model) evaluation.
  • Demonstrated ability to research and develop innovative methods for evaluating LLM-based, non-deterministic systems.
  • Working knowledge of Python or another programming language commonly used for testing and data analysis.
  • Strong analytical, research, and problem-solving abilities, with comfort operating in an emerging technology area where best practices are still evolving.
  • Good written and verbal communication skills.
Preferred Qualifications (Nice to Have)
  • Experience with DeepEval or similar LLM evaluation frameworks – a plus, not required.
  • Broader exposure to Generative AI, Machine Learning, or Natural Language Processing.
  • Experience evaluating AI-powered applications such as chatbots, copilots, or intelligent assistants.
  • Familiarity with prompt engineering concepts or major AI platforms (e.g., OpenAI, Anthropic, Google Gemini).
  • Experience in Financial Services or another regulated industry.
Success Profile

The ideal candidate is naturally curious, analytical, and comfortable working in emerging technology areas where best practices are still evolving. They can evaluate complex problems objectively, rapidly learn new concepts, and create practical approaches for measuring AI quality and performance. Most importantly, they possess the ability to research and develop innovative methods for LLM evaluation and translate those methods into scalable and repeatable testing practices.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Senior QA & AI Evaluation Scientist
Remote Senior QA & AI Evaluation Scientist

SDH Systems • United States

Remote
USD 120,000 - 155,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

On-site
USD 8,265 - 89,544
Secure computer and high-speed internet required
Automation QA w/ AI testing
Automation QA w/ AI testing

Compunnel, Inc. • New York (NY), Northern (KY)

On-site
USD 140,000 - 190,000
QA / Automation Engineer Agentic AI
QA / Automation Engineer Agentic AI

Compunnel, Inc. • Atlanta (GA), Northern (KY)

On-site
USD 110,000 - 160,000
QA Engineer – AI/ML
QA Engineer – AI/ML

Apptad Inc • Frisco (TX)

On-site
USD 120,000 - 160,000
QA Engineer - Agentic Systems
QA Engineer - Agentic Systems

Meet Life Sciences • New York (NY)

On-site
USD 110,000 - 170,000
AI QA Engineer
AI QA Engineer

Globant • North Carolina

On-site
USD 90,000 - 105,000
Generative AI Engineer (LLMs, MLOps, Remote)
Generative AI Engineer (LLMs, MLOps, Remote)

ClinDCast LLC • Warren Township (NJ)

On-site
USD 120,000 - 155,000
QA Engineer
QA Engineer

WebSenor Ltd • United States

On-site
USD 70,000 - 110,000
Lead QA Engineer
Lead QA Engineer

7Seventy • Northern (KY)

Hybrid
USD 140,000 - 165,000
Equity