AI Safety Engineer: Threat Models & Safe Systems

Aquent

New York (NY)

On-site

USD 99,000 - 107,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefit contributions
Retirement plan with match
Flexible spending accounts

Job summary

Skill connects the best professional talent with the world’s biggest brands. We are seeking a Safety Evaluations Engineer to develop threat models and design adversarial evaluations for AI systems, building reusable Python pipelines and LLM-based evaluation workflows.

You will translate findings into practical mitigations and partner with Engineering and Trust & Safety to embed evaluations into product development and monitoring loops, communicating results to diverse stakeholders.

Qualifications

  • Experience delivering safety evaluations or mitigations for real AI products with automation to scale processes.
  • Strong coding and practical data-analysis skills, with sufficient programming and SQL to judge code and queries independently.
  • Experience designing adversarial tests, evaluation datasets, taxonomies, rubrics, and metrics.

Responsibilities

  • Develop threat models and harm taxonomies for conversational, recommender, and tool-using AI systems.
  • Design and run single- and multi-turn adversarial evaluations with red teaming, automated attacks, synthetic data, and production data.
  • Build reusable Python evaluation pipelines, LLM-as-a-judge workflows, regression tests, dashboards, and curated golden datasets.
  • Validate evaluators against human labels and quantify coverage, reliability, false positives, false negatives, and safety–utility trade-offs.
  • Turn findings into practical mitigations including policy, prompts, context engineering, classifiers, data improvements, and system-level controls.
  • Work with Engineering and Trust & Safety to embed evaluations into product development and monitoring loops.
  • Communicate results clearly to both technical and non-technical stakeholders.
  • Full-time availability preferred; part-time arrangements may be considered.

Skills

Threat modeling
Adversarial evaluations
Python pipelines
Data analysis
SQL
Code reviews

Education

MSc or PhD in AI/ML

Tools

Python

Job description

Skill connects the best professional talent with the world’s biggest brands. We are seeking a Safety Evaluations Engineer to develop threat models and design adversarial evaluations for AI systems, building reusable Python pipelines and LLM-based evaluation workflows.

You will translate findings into practical mitigations and partner with Engineering and Trust & Safety to embed evaluations into product development and monitoring loops, communicating results to diverse stakeholders.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Safety Engineer
Applied AI Safety Engineer

Aquent • New York (NY)

On-site
USD 99,000 - 107,000
Health benefit contributions
Retirement plan with match
Flexible spending accounts
Senior Member of Technical Staff - Model Safety
Senior Member of Technical Staff - Model Safety

Xcede • San Francisco (CA)

On-site
USD 130,000 - 160,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • New York (NY)

On-site
USD 120,000 - 190,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 160,000
null
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • New York (NY)

On-site
USD 120,000 - 180,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Mercor • New York (NY)

On-site
USD 120,000 - 170,000
AI Safety Red Team Engineer
AI Safety Red Team Engineer

Crossing Hurdles • United States

On-site
USD 120,000 - 180,000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Obsidian • New York (NY)

On-site
USD 120,000 - 180,000
Remote AI Security Engineer — Threat Modeling & Defense
Remote AI Security Engineer — Threat Modeling & Defense

Bright Vision Technologies • Framingham (MA)

On-site
USD 100,000 - 150,000
Equal Opportunity Employer