AI Safety Researcher

CTI Clinical Trial and Consulting Services

Roma

In loco

EUR 60.000 - 90.000

Tempo pieno

3 giorni fa
Candidati tra i primi
Generatore di candidature

Trasforma questo ruolo in un colloquio — un curriculum e una lettera di presentazione personalizzati in base a ciò che cerca questo datore di lavoro.

Supera i filtri ATS

Descrizione del lavoro

Icaro Foundation in Rome seeks a mid- to senior-level researcher to advance safety research for frontier AI systems. You will review literature, design experiments, implement code, and contribute to papers and technical reports.

The role blends research and engineering: write Python code, extend tools, document configurations, and share reproducible results. Remote options may be considered for Europe or China, with in-person collaboration in Rome preferred.

Competenze

  • Experience with LLM evaluations, red-teaming, or benchmark development.
  • Strong technical or quantitative background from university study or projects.
  • Strong Python skills and familiarity with Git and debugging.
  • Experience with ML or LLMs through at least one project or research contribution.
  • Understanding of experimental reasoning, statistics, and uncertainty.
  • Ability to read technical papers and explain methods in English.
  • Strong interest in AI safety and urgency to reduce AI risks.
  • Willingness to revise views based on evidence.

Mansioni

  • Review literature, compare methods, and identify questions worth investigating.
  • Turn research questions into experimental protocols with baselines and evaluation criteria.
  • Implement and run experiments with frontier and open-weight models.
  • Analyze results and model traces, assess confounders and alternative explanations.
  • Contribute to research papers, benchmarks, and technical reports.
  • Develop the research codebase with Python, extend tools, and fix bugs.
  • Add tests and reproducible configurations for inspectability.
  • Document experiments, findings, and limitations.

Conoscenze

LLM evaluations
Red-teaming
Benchmark development
Python programming
Experimental design
Statistics basics
Critical reading
AI safety interest
Problem solving

Strumenti

Git

Descrizione del lavoro

The work

Icaro Foundation is an independent non-profit AI safety lab based in Rome. We study advanced AI systems: what they can do, how they fail, and how those findings can support developers and institutions responsible for their governance.

We see AI safety as one of the defining scientific and societal challenges of our time. As AI systems become more capable, autonomous, and widely deployed, understanding and reducing their risks is increasingly urgent. We are looking for people who are deeply interested in these questions and motivated to contribute through rigorous research.

You will help produce new research and develop the lab’s shared codebase and knowledge base, working closely with our researchers across the research process: reviewing literature, refining questions, implementing experiments, analysing results, and contributing to papers and technical reports.

Our research focuses particularly on agentic, multi-agent, and compositional safety: how risks emerge across extended interactions, tool use, and systems involving multiple AI agents. We also study testing awareness and evaluation validity, including whether models behave differently when they recognise that they are being evaluated.

Alongside our research, we evaluate frontier models for international model providers as independent third-party evaluators, using public and proprietary benchmarks and red-teaming environments.

Our public work includes:

  • Boiling the Frog, on multi-turn agentic safety;
  • Adversarial Humanities Benchmark, on the robustness of safety behaviour under stylistic reformulations;
  • research on LLM-to-LLM risks, multi-agent collusion, and interaction-level safety.

You can explore our research programme and papers to learn more.

What you would do

Your work will combine three closely connected areas.

Contribute to research

  • Review relevant literature, compare methods, and identify questions worth investigating.
  • Help turn research questions into experimental protocols, including baselines, controls, and clear evaluation criteria.
  • Implement and run experiments with frontier and open-weight models, including agentic and multi-agent environments.
  • Analyse results and model traces, investigate unexpected behaviour, and assess confounders and alternative explanations.
  • Contribute to research papers, benchmarks, technical reports, and presentations.

Develop the research codebase

  • Write and improve Python code for experiments, evaluations, data processing, and analysis.
  • Extend existing tools and environments, fix bugs, and participate in code review.
  • Add tests, documentation, and reproducible configurations so other researchers can inspect, rerun, and build on your work.

Build the lab’s knowledge base

  • Produce concise, source-grounded notes on papers, methods, benchmarks, and research questions.
  • Document experimental setups, findings, limitations, and negative results.
  • Organise and connect references, datasets, code, and research notes so the team can find relevant evidence and reuse previous work.

You may bring stronger skills in research or engineering. The role involves both writing code and reasoning carefully about evidence.

Who should apply

We welcome applications from master’s students, PhD students, recent graduates, and researchers at the beginning of their careers, including those who have recently completed a PhD.

Relevant experience may come from a thesis, academic research, independent experiments, open-source contributions, internships, or previous employment. We also welcome applicants from non-traditional backgrounds who can demonstrate strong research or engineering ability.

A completed PhD, previous AI safety employment, and published papers are not required. We care about the quality of your work, your contribution to it, and your ability to learn.

If you are currently studying, please tell us about your availability and how you would combine the role with your academic commitments.

What we are looking for

  • Experience with LLM evaluations, red-teaming, or benchmark development;
  • A solid technical or quantitative background, developed through university study, independent projects, or relevant work.
  • Good Python skills and familiarity with Git, debugging, and working with an existing codebase.
  • Practical experience with machine learning or LLMs through at least one substantive project or research contribution.
  • An understanding of basic experimental reasoning and statistics: comparing conditions, interpreting results, and recognising uncertainty and possible confounders.
  • The ability to read technical papers critically and explain methods, findings, and limitations clearly in English.
  • A strong interest in AI safety and a sense of urgency about understanding and reducing the risks posed by increasingly capable AI systems.
  • Intellectual curiosity, openness to criticism, and a willingness to revise your views in response to evidence.

Useful, not required

Experience with:

  • agentic or multi-agent systems;
  • statistical analysis or experimental replication;
  • software testing, containers, or reproducible research workflows;
  • literature reviews, research documentation, or open-source contributions;
  • Inspect AI, the open-source evaluation framework developed by the UK AI Security Institute and Meridian Labs, or comparable tools.

For an example of our research software, see the Adversarial Humanities Benchmark codebase, also listed in Inspect Evals as an externally maintained evaluation.

You do not need experience in all of these areas.

How we work

We are a small research team. You will work closely with experienced researchers and receive feedback on experimental design, code, analysis, and writing.

You will begin with clearly scoped contributions to ongoing projects and take on greater responsibility as your skills and familiarity with the work develop. We encourage everyone to ask questions, challenge assumptions, and propose ideas.

Existing evaluation infrastructure, technical support, and API budget are available. Contributions may become public papers, benchmarks, datasets, or tools where compatible with confidentiality obligations. Authorship and acknowledgement will reflect contributions.

We value work that others can understand and build on: clear reasoning, reliable code, well-documented experiments, and honest reporting of uncertainty.

  • Location: Flexible, with a preference for working in person with the team in Rome, Italy.
  • In-person collaboration: We particularly welcome applicants who are based in Rome or would be interested in relocating. We value regular in-person discussion, collaborative experimentation, and learning from one another.
  • Remote arrangements: May be considered for candidates based in Europe or China, with substantial overlap with European working hours.
  • Engagement: Contractor role.
  • Compensation: The specific compensation range will be shared during the first interview.

Referrals increase your chances of interviewing at Icaro Foundation by 2x

Find curated posts and insights for relevant topics all in one place.

Seniority level
  • Mid-Senior level
Employment type
  • Contract
Job function
  • Engineering and Information Technology
  • IT System Testing and Evaluation
Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Postdoctoral Research position on Adversarial Machine Learning & Formal Verification
Postdoctoral Research position on Adversarial Machine Learning & Formal Verification

AI4I • Collegno

In loco
EUR 30.000 - 55.000
Competitive salary and bonus incentives
Access to high-performance computing infrastructure
Opportunities for professional development
Research Engineer Position on Secure Agentic AI Systems
Research Engineer Position on Secure Agentic AI Systems

AI4I • Collegno

In loco
EUR 45.000 - 75.000
Competitive compensation
Full support for conference travel
Professional development opportunities
+2
AI Safety Researcher: Multi-Agent & Evaluation
AI Safety Researcher: Multi-Agent & Evaluation

CTI Clinical Trial and Consulting Services • Roma

In loco
EUR 60.000 - 90.000
Two Postdoctoral Positions in Robot Learning and Embodied AI
Two Postdoctoral Positions in Robot Learning and Embodied AI

The Italian Institute of Artificial Intelligence (AI4I) • Torino

In loco
EUR 42.000 - 56.000
International collaborations
Competitive compensation
Access to AI Foundry (Peano)
Two Postdoctoral Positions in Robot Learning and Embodied AI
Two Postdoctoral Positions in Robot Learning and Embodied AI

AI4I Foundation • Torino

In loco
EUR 30.000 - 55.000
Applied AI Architect, Industries Milan, Italy
Applied AI Architect, Industries Milan, Italy

Anthropic Limited • Milano

Ibrido
EUR 90.000 - 130.000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+1
Call for Founding Heads of R&D Laboratories
Call for Founding Heads of R&D Laboratories

The Italian Institute of Artificial Intelligence (AI4I) • Torino

In loco
EUR 70.000 - 120.000
Relocation support
Start-up grant
Tax exemption up to 90%
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Mercor • Roma

In loco
EUR 65.000 - 95.000
Senior Specialist in Administration, Finance and Accounts Payable Management
Senior Specialist in Administration, Finance and Accounts Payable Management

The Italian Institute of Artificial Intelligence (AI4I) • Torino

Ibrido
EUR 55.000 - 66.000
Modello ibrido di lavoro
Incentivi di relocation
Ambiente di lavoro a Torino
4th Call for Founding Heads of R&D Laboratories
4th Call for Founding Heads of R&D Laboratories

AI4I Foundation • Piemonte

Ibrido
EUR 70.000 - 120.000