LLM Evaluation Scenario Architect

Mindrift

Mississippi

Remote

USD 90,921 - 129,494

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive hourly rates
Flexible work hours
Valuable experience for portfolio

Job summary

A leading AI solutions company seeks a candidate for evaluating LLM-based agents. The role involves designing scenarios, creating test cases, and analyzing decision-making processes. A Bachelor's or Master's in relevant fields is required, alongside skills in data analysis and communication. The position offers flexible, remote work with rates up to $80/hour, allowing you to influence AI model developments while fitting around your current commitments.

Qualifications

  • Bachelor's and/or Master's in Computer Science, Software Engineering, AI or related fields.
  • Background in QA, software testing, or data analysis.
  • Good understanding of test design principles like reproducibility and edge cases.

Responsibilities

  • Design realistic evaluation scenarios for LLM-based agents.
  • Create structured test cases simulating human workflows.
  • Analyze agent logs and decision paths.

Skills

Analytical mindset
Attention to detail
Strong written communication skills in English
Understanding of test design principles
Curiosity about AI-generated content

Education

Bachelor's and/or Master's Degree in relevant fields

Tools

Python
JavaScript
JSON/YAML

Job description

A leading AI solutions company seeks a candidate for evaluating LLM-based agents. The role involves designing scenarios, creating test cases, and analyzing decision-making processes. A Bachelor's or Master's in relevant fields is required, alongside skills in data analysis and communication. The position offers flexible, remote work with rates up to $80/hour, allowing you to influence AI model developments while fitting around your current commitments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Agent Testing & Evaluation Scenario Architect
AI Agent Testing & Evaluation Scenario Architect

Mindrift • Houston (TX)

Remote
USD 60,000 - 80,000
Remote Applied AI Research Scientist (LLM & Evaluation)
Remote Applied AI Research Scientist (LLM & Evaluation)

Rex.zone • United States

Remote
USD 80,000 - 100,000
LLM Evaluation & Verification Lead
LLM Evaluation & Verification Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
LLM Model Response Evaluation
LLM Model Response Evaluation

Lifted, an Upwork Company™ • United States

On-site
AI Legal Researcher & Annotator for LLMs
AI Legal Researcher & Annotator for LLMs

Turing • Seattle (WA)

Remote
USD 125,000 - 150,000
Remote Part-Time Evaluation Scenario Writer for AI Testing
Remote Part-Time Evaluation Scenario Writer for AI Testing

Mindrift • Town of Texas (WI)

Remote
USD 83,000 - 110,000
Flexibility in work schedule
Competitive rates up to $80/hour
Opportunity to work on advanced AI projects
LLM Agent Evaluation Engineer — Build Intelligent Agents
LLM Agent Evaluation Engineer — Build Intelligent Agents

Bytedance • San Jose (CA)

On-site
USD 154,000 - 256,000
AI Legal Researcher & Annotator for LLMs
AI Legal Researcher & Annotator for LLMs

Turing • Los Angeles (CA)

Remote
USD 90,000 - 150,000
Remote AI/ML Research Engineer — LLM Training & Evaluation
Remote AI/ML Research Engineer — LLM Training & Evaluation

Rex.zone • United States

Remote
USD 80,000 - 100,000
Agent AI Engineer: LLM & Autonomous Systems
Agent AI Engineer: LLM & Autonomous Systems

TRM Labs Inc. • California (MO)

On-site
USD 200,000 - 275,000