Remote AI Evaluation Engineer: Safeguard & Scale

DeepRec.ai

Denver (CO)

Remote

USD 162,000 - 198,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

A mission-driven tech company is seeking an AI Evaluation Engineer to design and own evaluation systems that safeguard AI features. In this role, you will create frameworks and tools to ensure that AI is deployed safely and accurately. The ideal candidate has a strong software engineering background and experience with OpenAI API or similar LLM tools. Join a team that empowers professionals to deliver critical support through AI capabilities while prioritizing safety and innovation.

Qualifications

  • Experience with TypeScript is a plus.
  • Practical knowledge of function calling and LLM grading.
  • Ability to validate data quality and performance.

Responsibilities

  • Design frameworks for evaluation outputs.
  • Build data pipelines and integrate with CI.
  • Conduct red team assessments of AI systems.
  • Establish model versioning and observability.
  • Deliver tooling and dashboards for engineers.

Skills

Strong software engineering background
Deep experience with OpenAI API or similar LLM ecosystems
Practical knowledge of prompting and eval techniques
Familiarity with statistical analysis
Experience with observability or data science tooling

Job description

A mission-driven tech company is seeking an AI Evaluation Engineer to design and own evaluation systems that safeguard AI features. In this role, you will create frameworks and tools to ensure that AI is deployed safely and accurately. The ideal candidate has a strong software engineering background and experience with OpenAI API or similar LLM tools. Join a team that empowers professionals to deliver critical support through AI capabilities while prioritizing safety and innovation.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

On-site
USD 162,000 - 198,000
AI Evaluation Lead: Real-World Systems Benchmarking
AI Evaluation Lead: Real-World Systems Benchmarking

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
Remote AI Data Engineer: LLM Evaluation & Validation
Remote AI Data Engineer: LLM Evaluation & Validation

Veeva Systems • Portland (ME)

Hybrid
USD 85,000 - 225,000
Medical, dental, and vision insurance
Flexible PTO
Retirement programs
+1
AI Risk & Fraud Evaluation Engineer
AI Risk & Fraud Evaluation Engineer

Variance • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive salary
Platinum-level medical, dental, and vision insurance
Unlimited PTO
+2
AI Evaluation Scientist: Trust & Safety Leader
AI Evaluation Scientist: Trust & Safety Leader

Steampunk • McLean (VA)

Hybrid
USD 105,000 - 145,000
AI Security Engineer — Secure AI-Driven Apps
AI Security Engineer — Secure AI-Driven Apps

Fortinet • Sunnyvale (CA)

On-site
USD 160,000 - 220,000
Medical insurance
Dental insurance
401(k)
+3
Remote AI Safety & Biosecurity Software Engineer
Remote AI Safety & Biosecurity Software Engineer

SecureBio, LLC • Boston (MA)

Remote
USD 70,000 - 115,000
Unlimited paid time off
Flexible work hours
Conference sponsorships
+1
AI Evaluation Scientist — Metrics, Safety & Trust
AI Evaluation Scientist — Metrics, Safety & Trust

UNAVAILABLE • McLean (VA)

On-site
USD 120,000 - 150,000
Delivery Engineer: AI Safety & Enterprise Evaluations
Delivery Engineer: AI Safety & Enterprise Evaluations

Artificial Intelligence Underwriting Company • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive salary
Equity
Relocation to San Francisco
+1
Senior AI Evaluation Engineer — Metrics & Data Pipelines
Senior AI Evaluation Engineer — Metrics & Data Pipelines

Sentry • San Francisco (CA)

Hybrid
USD 240,000 - 280,000
Equity grants
Paid time off
Group health insurance coverage