AI Evaluation Engineer

DeepRec.ai

Denver (CO)

Remote

USD 162,000 - 198,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

A mission-driven tech company is seeking an AI Evaluation Engineer to design and own evaluation systems that safeguard AI features. In this role, you will create frameworks and tools to ensure that AI is deployed safely and accurately. The ideal candidate has a strong software engineering background and experience with OpenAI API or similar LLM tools. Join a team that empowers professionals to deliver critical support through AI capabilities while prioritizing safety and innovation.

Qualifications

  • Experience with TypeScript is a plus.
  • Practical knowledge of function calling and LLM grading.
  • Ability to validate data quality and performance.

Responsibilities

  • Design frameworks for evaluation outputs.
  • Build data pipelines and integrate with CI.
  • Conduct red team assessments of AI systems.
  • Establish model versioning and observability.
  • Deliver tooling and dashboards for engineers.

Skills

Strong software engineering background
Deep experience with OpenAI API or similar LLM ecosystems
Practical knowledge of prompting and eval techniques
Familiarity with statistical analysis
Experience with observability or data science tooling

Job description

US Recruitment Consultant: Guiding GenAI professionals towards their dream careers

AI Evaluation Engineer
$180,000
Remote (US-based)

Are you passionate about shaping how AI is deployed safely, reliably, and at scale? This is a rare opportunity to join a mission‑driven tech company as their first AI Evaluation Engineer, a foundational role where you’ll design, build, and own the evaluation systems that safeguard every AI‑powered feature before it reaches the real world.

This organization builds AI‑enabled products that directly helps governments, nonprofits, and agencies deliver financial support to people who need it most. As AI capabilities race forward, ensuring these systems are safe, accurate, and resilient is critical. That’s where you come in.

You won’t just be testing models, you’ll be creating the frameworks, pipelines, and guardrails that make advanced LLM features safe to ship. You’ll collaborate with engineers, PMs, and AI safety experts to stress test boundaries, uncover weaknesses, and design scalable evaluation systems that protect end users while enabling rapid innovation.

What You’ll Do
  • Own the evaluation stack – design frameworks that define “good,” “risky,” and “catastrophic” outputs.
  • Automate at scale – build data pipelines, LLM judges, and integrate with CI to block unsafe releases.
  • Stress testing – red team AI systems with challenge prompts to expose brittleness, bias, or jailbreaks.
  • Track and monitor – establish model/prompt versioning, build observability, and create incident response playbooks.
  • Empower others – deliver tooling, APIs, and dashboards that put eval into every engineer’s workflow.
Requirements
  • Strong software engineering background (TypeScript a plus)
  • Deep experience with OpenAI API or similar LLM ecosystems
  • Practical knowledge of prompting, function calling, and eval techniques (e.g. LLM grading, moderation APIs)
  • Familiarity with statistical analysis and validating data quality/performance
  • Bonus: experience with observability, monitoring, or data science tooling
Seniority level
  • Not Applicable
Employment type
  • Full-time
Job function
  • Information Technology
  • Technology, Information and Media, Information Services, and Software Development
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Backend Software Engineer (Evals)
Backend Software Engineer (Evals)

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 385,000
AI Evaluation Infrastructure Consultant
AI Evaluation Infrastructure Consultant

Landing Point • Village of Pelham (NY)

On-site
USD 152,000 - 179,000
Engineering Manager – Evaluation & Observability
Engineering Manager – Evaluation & Observability

AgentsFlow • Northern (KY)

On-site
USD 170,000 - 190,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

On-site
USD 8,265 - 89,544
Secure computer and high-speed internet required
LLM Evaluation Engineering Lead
LLM Evaluation Engineering Lead

DeepRec.ai • Redwood City (CA)

On-site
USD 180,000 - 240,000
High autonomy
Strong technical peers
Meaningful equity
AI Engineer
AI Engineer

7Seventy • Northern (KY)

On-site
USD 130,000 - 150,000
Remote work nationwide
Health benefits
Competitive compensation package
+2
Senior AI Engineer
Senior AI Engineer

arosplatforms | AI Consulting & Services • North Township (IN)

On-site
USD 120,000 - 170,000
Backend Software Engineer (Evals)
Backend Software Engineer (Evals)

OpenAI • Seattle (WA)

On-site
USD 230,000 - 385,000
AI Evaluation Engineer – Reinforcement Learning & Agents
AI Evaluation Engineer – Reinforcement Learning & Agents

MaxIT Consulting - Max Corporate Group • San Francisco (CA)

On-site
USD 140,000 - 210,000
Staff AI Engineer
Staff AI Engineer

Harnham • San Francisco (CA)

On-site
USD 180,000 - 240,000