AI Safety Research Scientist: Trace & Defuse Agent Failures

Whitecircle

Paris

Hybride

EUR 90 000 - 140 000

Plein temps

Il y a 7 jours
Soyez parmi les premiers à postuler
Générateur de candidature

N’envoyez pas un CV générique — générez un CV et une lettre de motivation adaptés à ce poste précis.

Passez les filtres ATS

Avantages offerts par ce poste

Paid time off
Hybrid Paris w relocation
Medical insurance
Equity package
Hardware & tools
Subscriptions for AI agents

Résumé du poste

White Circle is seeking a research scientist to study how LLM agents fail in the wild, eliciting deception and unsafe behaviours, and building a deeper understanding of misalignment in realistic user scenarios.

You will own end-to-end research from unclear questions to falsifiable experiments, develop audit agents at scale, and publish findings to inform guardrails and product safety. Hybrid Paris-based role with relocation support potential.

Qualifications

  • Track record of empirical research in agent behaviour, model evaluation, alignment, or closely adjacent area.
  • Able to build a research MVP involving fine-tuning, agent inference, and evals without platform team support.
  • Experience designing experiments under real conditions: isolating failures, calibrating baselines, distinguishing signal from artifacts.

Responsabilités

  • Own research projects end to end from unclear concerns to falsifiable experiments and defendable results.
  • Develop automated audit agents that discover and characterize model behavior at scale.
  • Study misalignment and bias as users interact with agents and turn findings into evals for products.
  • Pressure-test frontier agents in realistic, high-stakes scenarios to find failure points before customers.
  • Run white-box and black-box investigations to understand how AI models fail.
  • Publish learnings as public blog posts and conference papers and feed back into guardrails.

Connaissances

Empirical research
ML engineering
Experimental design
AI safety
Frontier models

Formation

MSc or PhD in ML/CS/ cognitive science

Outils

Fine-tuning
Agent inference
Eval frameworks

Description du poste

White Circle is seeking a research scientist to study how LLM agents fail in the wild, eliciting deception and unsafe behaviours, and building a deeper understanding of misalignment in realistic user scenarios.

You will own end-to-end research from unclear questions to falsifiable experiments, develop audit agents at scale, and publish findings to inform guardrails and product safety. Hybrid Paris-based role with relocation support potential.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

AI Safety Research Scientist: Agent Misbehavior & Alignment
AI Safety Research Scientist: Agent Misbehavior & Alignment

Aisafety • Paris

Hybride
EUR 70 000 - 110 000
Paid time off
Relocation package
Medical insurance (France)
+3
Research Scientist, Agentic Systems & AI Safety
Research Scientist, Agentic Systems & AI Safety

Whitecircle • Paris

Hybride
EUR 90 000 - 150 000
Paid time off (local regulations)
Relocation package for Paris or London
Equity package
+2
Research Scientist (AI Behaviours)
Research Scientist (AI Behaviours)

White Circle • Paris

Hybride
EUR 131 000 - 220 000
Relocation package
Comprehensive medical insurance
All hardware, tools & services
+2
Research Scientist, AI Safety & Multi-Agent Environments
Research Scientist, AI Safety & Multi-Agent Environments

White Circle • Paris

Hybride
EUR 131 000 - 219 000
Paid time off
Comprehensive medical insurance
Team off-sites twice a year
Research Scientist/Engineer (Agentic Systems)
Research Scientist/Engineer (Agentic Systems)

White Circle • Paris

Hybride
EUR 131 000 - 219 000
Paid time off
Comprehensive medical insurance
Team off-sites twice a year
Research Scientist/Engineer (Agentic Systems)
Research Scientist/Engineer (Agentic Systems)

Whitecircle • Paris

Hybride
EUR 90 000 - 150 000
Paid time off (local regulations)
Relocation package for Paris or London
Equity package
+2
Adversarial AI Environment Research Engineer
Adversarial AI Environment Research Engineer

Visa Hunt • Paris

Hybride
EUR 90 000 - 150 000
Equity package
Medical insurance (France)
Relocation support
+1
Research Scientist (AI Behaviours)
Research Scientist (AI Behaviours)

Whitecircle • Paris

Hybride
EUR 90 000 - 140 000
Paid time off
Hybrid Paris w relocation
Medical insurance
+3
Research Scientist (AI Behaviours)
Research Scientist (AI Behaviours)

Aisafety • Paris

Hybride
EUR 70 000 - 110 000
Paid time off
Relocation package
Medical insurance (France)
+3
Research Engineer (Evals)
Research Engineer (Evals)

White Circle • Paris

Hybride
EUR 131 000 - 219 000
Paid time off
Comprehensive medical insurance
Team off-sites twice a year