AI Safety & Evaluation Engineer

DeWinter Group

Campbell (CA)

Remote

USD 68,880 - 241,080

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI solutions firm is seeking an AI Safety and Evaluations Engineer for a 12-month contract, focusing on designing evaluation frameworks to ensure AI models are bias-free and compliant. The candidate will create automated datasets and develop specific metrics in RAG-based systems. With a requirement of 3+ years in AI Research or Quality Engineering, strong communication skills, and autonomy, this role is integral to maintaining the safety of AI models. This is a remote opportunity with a pay range of $50/hr to $175/hr.

Qualifications

  • 3+ years of experience in AI Research or Quality Engineering.
  • Deep expertise in model evaluation techniques and NLP metrics.
  • Demonstrated ability to work autonomously and manage time effectively.
  • Experience with Python, data analysis tools, and LLM-as-a-Judge frameworks.
  • Strong communication skills for team updates.

Responsibilities

  • Design and build evaluation frameworks for model bias.
  • Create automated datasets to benchmark models before production.
  • Develop metrics for 'Grounding' and 'Faithfulness' in RAG-based systems.
  • Build monitoring tools for harmful AI outputs.
  • Partner with legal and ethics teams for safety constraints.

Skills

AI Research
Quality Engineering
Model evaluation techniques
NLP metrics (ROUGE, BLEU, BERTScore)
Python
Data analysis tools
Self-motivated
Communication skills

Job description

Title: AI Safety and Evaluations Engineer
Job Type: Contract
Contract Length: 12 Months
Pay Range: $50/hr – $175/hr
Start Date: ASAP
Location: Remote

About the Opportunity: Our client, a leader in AI testing and Generative AI solutions, is looking for a skilled AI Safety and Evaluations Engineer to join their team for a 12-month engagement. This project involves designing and building rigorous evaluation frameworks to measure model bias, hallucinations, and toxicity, ensuring models are safe and compliant before deployment. This is a high-impact role that requires a self-motivated professional who can hit the ground running and deliver results quickly.

Key Responsibilities & Deliverables
  • Designing and building rigorous evaluation frameworks to measure model bias, hallucinations, and toxicity.
  • Creating automated "Eval" datasets to benchmark new models before they are promoted to production.
  • Developing metrics for "Grounding" and "Faithfulness" in RAG-based systems.
  • Building monitoring tools that flag harmful or non-compliant AI outputs in real-time.
  • Partnering with legal and ethics teams to translate policy into technical safety constraints.
Required Skills & Experience
  • 3+ years of experience in AI Research or Quality Engineering.
  • Deep expertise in model evaluation techniques and NLP metrics (ROUGE, BLEU, BERTScore). This isn't a learning role—you need to be a subject matter expert.
  • Demonstrated ability to work autonomously and manage your own time effectively to meet project goals.
  • Experience with Python, data analysis tools, and LLM-as-a-Judge frameworks.
  • Strong communication skills to provide clear and concise status updates to the project team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
Remote AI Safety & Evaluations Engineer
Remote AI Safety & Evaluations Engineer

DeWinter Group • Campbell (CA)

Remote
USD 68,880 - 241,080
AI Safety Specialist - Fully Remote | Upto $70/hr
AI Safety Specialist - Fully Remote | Upto $70/hr

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 180,000
AI Evaluation Scientist
AI Evaluation Scientist

Steampunk, Inc. • McLean (VA)

On-site
USD 140,000 - 210,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Mercor • New York (NY)

On-site
USD 120,000 - 170,000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Obsidian • New York (NY)

On-site
USD 120,000 - 180,000
AI Safety Practitioner
AI Safety Practitioner

Great Value Hiring • United States

On-site