AI Safety Research Scientist: Agent Misbehavior & Alignment

Aisafety

Paris

Hybride

EUR 70 000 - 110 000

Plein temps

14 jours+

Recevez plus de réponses des employeurs

Envoyez un CV adapté au poste en quelques minutes.

Avantages offerts par ce poste

Paid time off
Relocation package
Medical insurance (France)
Hardware & tools provided
Subscriptions for AI agents and IDEs
Team off-sites

Résumé du poste

White Circle is seeking a research scientist to study how LLM agents fail in real-world settings, eliciting deception and unsafe behaviour in concrete experiments. You will build understanding of how agents misbehave with realistic user interactions and develop evals for our products.

You will publish findings and contribute to internal guardrails, working in a small, focused team with opportunities to ship notable safety improvements.

Qualifications

  • A track record of empirical research in agent behaviour, model evaluation, alignment, or adjacent area.
  • Strong ML engineering: independently build a research MVP involving fine-tuning, agent inference, and evals.
  • Experimental design under real conditions: isolating agent failure modes, calibrating judges and baselines.
  • You can take a vague behavioural question and define the experiment that answers it, then run it fast.
  • An AI power-user—fluent with frontier models and coding agents in daily work.

Responsabilités

  • Own research projects end to end—from unclear concerns to falsifiable experiments and defend results.
  • Develop automated audit agents that discover and characterize suspect model behaviour at scale.
  • Study misalignment and bias when real users interact with agents and turn findings into evals.
  • Pressure-test frontier agents in realistic, high-stakes scenarios to find where they break.
  • Run white-box and black-box investigations to understand how AI models fail.
  • Publish what you learn as public blog posts and conference papers, feeding back into guardrails.

Connaissances

Empirical research
ML engineering
Experimental design
Frontier models
AI power-user

Formation

MSc or PhD in ML/CS/cognitive science/comp neuroscience/physics

Outils

NLAs
SAEs

Description du poste

White Circle is seeking a research scientist to study how LLM agents fail in real-world settings, eliciting deception and unsafe behaviour in concrete experiments. You will build understanding of how agents misbehave with realistic user interactions and develop evals for our products.

You will publish findings and contribute to internal guardrails, working in a small, focused team with opportunities to ship notable safety improvements.

Obtenez votre examen gratuit et confidentiel de votre CV.
ou faites glisser et déposez votre fichier ici.
Similar jobs

Postes similaires à comparer

AI Safety Research Scientist — Adversarial Environments
AI Safety Research Scientist — Adversarial Environments

White Circle • Paris

Hybride
EUR 60 000 - 90 000
Paid time off
Comprehensive medical insurance
Covered subscriptions for AI tools
+1
Research Scientist, AI Safety & Behavior Evaluation
Research Scientist, AI Safety & Behavior Evaluation

Slope • Paris

Hybride
EUR 60 000 - 80 000
Comprehensive medical insurance
Paid time off
Relocation package available
+1
Research Scientist – AI Behaviours
Research Scientist – AI Behaviours

Slope • Paris

Hybride
EUR 60 000 - 80 000
Comprehensive medical insurance
Paid time off
Relocation package available
+1
Research Scientist (AI Behaviours)
Research Scientist (AI Behaviours)

White Circle • Paris

Hybride
EUR 131 000 - 220 000
Relocation package
Comprehensive medical insurance
All hardware, tools & services
+2
Research Scientist/Engineer – Agentic Systems
Research Scientist/Engineer – Agentic Systems

White Circle • Paris

Hybride
EUR 60 000 - 90 000
Paid time off
Comprehensive medical insurance
Covered subscriptions for AI tools
+1
Research Scientist, AI Safety & Multi-Agent Environments
Research Scientist, AI Safety & Multi-Agent Environments

White Circle • Paris

Hybride
EUR 131 000 - 219 000
Paid time off
Comprehensive medical insurance
Team off-sites twice a year
Adversarial AI Environment Research Engineer
Adversarial AI Environment Research Engineer

Visa Hunt • Paris

Hybride
EUR 90 000 - 150 000
Equity package
Medical insurance (France)
Relocation support
+1
Research Scientist/Engineer (Agentic Systems)
Research Scientist/Engineer (Agentic Systems)

White Circle • Paris

Hybride
EUR 131 000 - 219 000
Paid time off
Comprehensive medical insurance
Team off-sites twice a year
Research Scientist (AI Behaviours)
Research Scientist (AI Behaviours)

Aisafety • Paris

Hybride
EUR 70 000 - 110 000
Paid time off
Relocation package
Medical insurance (France)
+3
Research Engineer - AI Safety Benchmarks
Research Engineer - AI Safety Benchmarks

Visa Hunt • Paris

Hybride
EUR 85 000 - 125 000
Paid time off
Hybrid Paris work with relocation
France medical insurance
+3