Applied AI Safety Engineer

Aquent

New York (NY)

On-site

USD 99,000 - 107,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefit contributions
Retirement plan with match
Flexible spending accounts

Job summary

Skill connects the best professional talent with the world’s biggest brands. We are seeking a Safety Evaluations Engineer to develop threat models and design adversarial evaluations for AI systems, building reusable Python pipelines and LLM-based evaluation workflows.

You will translate findings into practical mitigations and partner with Engineering and Trust & Safety to embed evaluations into product development and monitoring loops, communicating results to diverse stakeholders.

Qualifications

  • Experience delivering safety evaluations or mitigations for real AI products with automation to scale processes.
  • Strong coding and practical data-analysis skills, with sufficient programming and SQL to judge code and queries independently.
  • Experience designing adversarial tests, evaluation datasets, taxonomies, rubrics, and metrics.

Responsibilities

  • Develop threat models and harm taxonomies for conversational, recommender, and tool-using AI systems.
  • Design and run single- and multi-turn adversarial evaluations with red teaming, automated attacks, synthetic data, and production data.
  • Build reusable Python evaluation pipelines, LLM-as-a-judge workflows, regression tests, dashboards, and curated golden datasets.
  • Validate evaluators against human labels and quantify coverage, reliability, false positives, false negatives, and safety–utility trade-offs.
  • Turn findings into practical mitigations including policy, prompts, context engineering, classifiers, data improvements, and system-level controls.
  • Work with Engineering and Trust & Safety to embed evaluations into product development and monitoring loops.
  • Communicate results clearly to both technical and non-technical stakeholders.
  • Full-time availability preferred; part-time arrangements may be considered.

Skills

Threat modeling
Adversarial evaluations
Python pipelines
Data analysis
SQL
Code reviews

Education

MSc or PhD in AI/ML

Tools

Python

Job description

What You’ll Do
  • Develop product-specific threat models and harm taxonomies for conversational, recommender, and tool-using AI systems.
  • Design and run single- and multi-turn adversarial evaluations, combining expert red teaming, automated attack generation, synthetic data, and sampled production data.
  • Build reusable Python evaluation pipelines, LLM-as-a-judge workflows, regression tests, dashboards, and curated golden datasets.
  • Validate evaluators against human labels and quantify coverage, judge reliability, false positives, false negatives, and safety–utility trade-offs.
  • Turn findings into practical mitigations, including policy and prompt changes, context engineering, classifiers, data improvements, preference tuning, and system-level controls.
  • Work directly with Engineering and Trust & Safety to embed evaluations into product-development and monitoring loops.
  • Communicate results clearly to technical and non-technical stakeholders.
  • Full-time availability is preferred, although part-time arrangements may be considered.
Who You Are
  • You have personally delivered safety evaluations or mitigations for a real AI or machine-learning product, using automation to scale processes.
  • You have strong coding agent and practical data-analysis skills, with sufficient programming and SQL to independently judge code and queries.
  • You have experience designing adversarial tests, evaluation datasets, taxonomies, rubrics, and metrics.
  • You can work autonomously when the risk, success criteria, and methodology are initially unclear.
  • You have experience working across research, engineering, product, policy, or Trust & Safety.
  • You communicate clearly in writing and have a record of turning research findings into action.
It’s a Plus If You Have
  • Experience evaluating multi-turn or tool-using agents.
  • Experience calibrating LLM judges or building human-in-the-loop evaluations.
  • Experience with multilingual or multimodal evaluation.
  • Experience with preference tuning or other model-alignment techniques.
  • Experience with multilingual or multimodal evaluation.
  • An MSc or PhD in an AI/ML-related field.

The target hiring compensation range for this role is $72.00/hr to $78.00/hr.

About Skill:

Skill connects the best professional, IT, engineering, financial and administrative talent with the world’s biggest brands. Our eligible talent get access to benefits such as health benefit contributions, retirement plans with match and flexible spending accounts.

Skill is an equal-opportunity employer. We evaluate qualified applicants without regard to age, race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, and other legally protected characteristics. We’re about creating an inclusive environment-one where different backgrounds, experiences, and perspectives are valued, and everyone can contribute, grow their careers, and thrive.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer
AI Evaluation Engineer

DeepRec.ai • Denver (CO)

Remote
USD 180,000
Applied AI Engineer
Applied AI Engineer

Phase2 Technology • Fort Belvoir (VA)

On-site
USD 77,000 - 176,000
Software Engineer, Infrastructure & Platform
Software Engineer, Infrastructure & Platform

10a Labs • United States

Remote
USD 110,000 - 160,000
Fully remote, U.S.-based
Performance-based annual bonus
Professional development support (con-
+1
AI Safety & Evaluation Engineer
AI Safety & Evaluation Engineer

DeWinter Group • Campbell (CA)

Remote
USD 68,880 - 241,080
AI Safety Specialist - Fully Remote | Upto $70/hr
AI Safety Specialist - Fully Remote | Upto $70/hr

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 180,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required
AI Safety Engineer: Threat Models & Safe Systems
AI Safety Engineer: Threat Models & Safe Systems

Aquent • New York (NY)

On-site
USD 99,000 - 107,000
Health benefit contributions
Retirement plan with match
Flexible spending accounts
AI Safety Expert (Seattle or Boston)
AI Safety Expert (Seattle or Boston)

Portland Seed Fund • Seattle (WA)

On-site
USD 55,000 - 103,000
Safeguards Enforcement Analyst, Safety Evaluations
Safeguards Enforcement Analyst, Safety Evaluations

United States Digital Space LLC • United States

Hybrid
USD 230,000 - 270,000
AI Safety Practitioner
AI Safety Practitioner

Great Value Hiring • United States

On-site