AI Safety Data Scientist

Aquent

New York (NY)

On-site

USD 120,000 - 160,000

Full time

12 days ago
Application generator

Get a reply from this recruiter — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Aquent is seeking an experienced Data Scientist to measure and monitor risks in conversational and agentic AI products. You will investigate emerging safety risks in production and develop scalable measurement, evaluation, and monitoring systems for policy and product improvements.

You will collaborate with Product, Engineering, and Trust & Safety teams, translate insights into mitigations, and communicate findings to stakeholders with clarity and impact.

Qualifications

  • Experience delivering safety evaluations for AI products.
  • Ability to design evaluation datasets and rubrics.
  • Strong written and verbal communication for stakeholder uptake.
  • Comfort structuring ambiguous problems and defining success criteria.

Responsibilities

  • Develop and scale risk monitoring and safety measurement systems across AI features.
  • Design safety metrics and evaluation frameworks for production systems (false positives/negatives, safety risk prevalence).
  • Inspect, debug, and use Python data analysis scripts and SQL pipelines (LLM tools optional).
  • Calibrate LLM-evaluation workflows, label datasets, and build dashboards for leadership.
  • Collaborate with Product, Engineering, and Trust & Safety to translate insights into mitigations and policies.
  • Communicate analytical findings clearly to technical and non-technical stakeholders.

Skills

Python
SQL
Data analysis
AI safety
Risk metrics

Tools

Claude
Codex

Job description

Location: Remote (US – East Coast / EST hours preferred)

Role Summary:

We are looking for an experienced Data Scientist to help us measure and monitor risks in conversational and agentic AI products. Reporting directly to the Data Science Lead, you will investigate emerging safety risks in production and turn them into scalable measurement, evaluation, and monitoring systems that inform policy, model, and product improvements.

What You’ll Do:
  • Develop and scale risk monitoring and safety measurement systems across conversational, recommender, and tool-using AI features.
  • Design safety metrics and evaluation frameworks for production systems (including false positive/negative rates, safety risk prevalence, and decision rubrics).
  • Inspect, debug, and utilize Python data analysis scripts and SQL pipelines (leveraging LLM coding tools such as Claude or Codex to expedite analytical workflows).
  • Calibrate LLM-as-a-judge evaluation workflows, label evaluation datasets, and build reporting dashboards/visualizations for leadership and product stakeholders.
  • Work cross-functionally with Product, Engineering, and Trust & Safety teams to translate real-world production insights into safety mitigations and updated safety policies.
  • Communicate analytical findings and measurement metrics clearly to both technical and non-technical stakeholders.
Who You Are:
  • You have personally delivered safety evaluations, risk metrics, or mitigations for a real AI or machine-learning product.
  • You have strong Python and SQL data analysis skills, with the technical ability to independently inspect, critique, and debug pipeline code and queries.
  • You have hands-on experience designing evaluation datasets, rubrics, and measurement metrics.
  • You are comfortable structuring ambiguous problems and defining success criteria, methodologies, and tradeoffs with stakeholders.
  • You have experience working cross-functionally across multiple domains, including research, engineering, product, policy, or Trust & Safety.
  • You communicate clearly in writing and have a proven record of turning analytical research findings into clear product/policy action.
  • Minimum 2 years of direct AI safety data science or safety evaluation experience, OR 3 to 5+ years of broader Data Science experience if safety experience is foundational or project-based.
It’s a Plus If You Have:
  • Experience calibrating LLM judges or building human-in-the-loop evaluations.
  • Experience evaluating multi-turn or tool-using agentic systems.
  • Experience with multilingual evaluation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Safety Data Scientist (Contract)
Remote AI Safety Data Scientist (Contract)

Skill • New York (NY)

On-site
USD 105,000 - 113,000
W2 benefits
401k matching
AI Safety Data Scientist — Remote Risk & Evaluation
AI Safety Data Scientist — Remote Risk & Evaluation

Aquent • New York (NY)

On-site
USD 120,000 - 160,000
Data Scientist, Meta Superintelligence Labs (Safety)
Data Scientist, Meta Superintelligence Labs (Safety)

Meta • Menlo Park (CA)

On-site
USD 180,000 - 230,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 160,000
null
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • New York (NY)

On-site
USD 120,000 - 180,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • New York (NY)

On-site
USD 120,000 - 190,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Applied AI Safety Engineer
Applied AI Safety Engineer

Aquent • New York (NY)

On-site
USD 99,000 - 107,000
Health benefit contributions
Retirement plan with match
Flexible spending accounts
AI Safety Practitioner
AI Safety Practitioner

Dorado • United States

Remote
USD 110,000 - 170,000
AI Safety Expert (Seattle or Boston)
AI Safety Expert (Seattle or Boston)

mpathic • Seattle (WA)

On-site
USD 70,000 - 110,000