AI Research Engineer

nineDots.io

Dublin

On-site

EUR 120,000 - 180,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

nineDots.io in Dublin is seeking an AI Research Engineer focused on frontier models and benchmarking. You will design and run rigorous evaluation protocols across multiple state-of-the-art LLMs and analyze model behavior.

The role values deep research credentials, publications or demonstrated benchmark work, and the ability to build evaluation systems and datasets. You’ll work on post-training and reinforcement learning experiments and compare model strengths across several families.

Qualifications

  • Deep expertise in frontier LLM benchmarks and model evaluation.
  • Publications or demonstrable benchmark work.
  • Experience building evaluation systems and datasets.

Responsibilities

  • Design benchmarks for frontier LLMs.
  • Compare model performance across families such as Claude, GPT, Gemini, Llama.
  • Develop datasets to expose strengths and weaknesses.
  • Conduct adversarial testing and reliability analyses.
  • Work on post-training and reinforcement learning experiments.

Skills

LLM benchmarking
Adversarial testing
AI safety research
Post-training RL
Synthetic data generation
Evaluation frameworks
Model evaluation
Multimodel analysis

Education

PhD in a relevant area
Publications at conferences (NeurIPS/ICML/ICLR)

Job description

AI Research Engineer – Frontier Models & Benchmarking

Dublin | 5 Days Onsite

We’re working with a well-funded AI company building training and evaluation environments for some of the world’s leading AI labs.

They’re growing a small research team in Dublin and are looking for people who are genuinely deep into frontier AI models.

This is a deliberately specialised role.

If your LLM experience is mainly building RAG applications, chatbots, agent orchestration or integrating models into existing products, this probably isn’t the right role for you.

They’re looking for people who study the models themselves.

You’ll be working on problems like:

  • Designing and building benchmarks for frontier LLMs
  • Comparing how different models perform and behave
  • Building difficult tasks and datasets to expose model strengths and weaknesses
  • Adversarially testing models and finding where they break
  • Designing rigorous evaluation methodologies
  • Working with synthetic training and evaluation data
  • Investigating model reasoning, reliability and unexpected behaviour
  • Working on post-training and reinforcement learning
  • Running experiments across multiple frontier models

The people we particularly want to hear from have experience in one or more of:

  • Building or contributing to recognised/public LLM benchmarks
  • LLM benchmarking or model capability evaluation
  • Adversarial LLM testing or red teaming
  • AI safety / model safety research
  • LLM post-training or reinforcement learning
  • Synthetic data generation for model training/evaluation
  • Designing evaluation frameworks, metrics or model judges
  • Research into LLM behaviour and failure modes
  • Working deeply across multiple models such as Claude, GPT, Gemini, Llama or DeepSeek

Strong research credentials are highly valued. That could mean a PhD in a relevant area, significant research experience, publications at conferences such as NeurIPS, ICML or ICLR, or demonstrably strong work building benchmarks and evaluation systems.

Most importantly, we’re looking for people who have actually done this work, rather than simply used the terminology.

If your experience is predominantly:

  • RAG
  • LangChain
  • Vector databases
  • Prompt engineering
  • Building chatbots
  • Agent orchestration
  • Calling LLM APIs
  • Adding GenAI features to existing products

...without deeper model evaluation, benchmarking or research experience, this role is unlikely to be a fit.

This is also 5 days per week onsite in Dublin city centre. It’s a small, ambitious team that moves quickly and works hard, so it’ll suit someone who actively wants that kind of start-up environment rather than a traditional 9-to-5.

If you’re the sort of person who sees a new frontier model released and immediately wants to test it, break it, compare it and understand why it behaves differently from the others, we’d like to hear from you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Frontier LLM Benchmark Research Engineer
Frontier LLM Benchmark Research Engineer

nineDots.io • Dublin

On-site
EUR 120,000 - 180,000
Senior AI Engineer (m/f/d)
Senior AI Engineer (m/f/d)

Jobgether • Ireland

On-site
EUR 120,000 - 180,000
Home-office budget
Learning budget
AI Applied Researcher
AI Applied Researcher

Accenture UK & Ireland • Dublin

On-site
EUR 90,000 - 130,000
Pension
Private health insurance
Gym membership discount
+1
AI Engineer
AI Engineer

Ingenium • Limerick

Hybrid
EUR 90,000 - 130,000
Performance bonus
Company pension
25 days leave
+2
Senior Applied AI Researcher (Dublin, CA)
Senior Applied AI Researcher (Dublin, CA)

Articul8 • Dublin

On-site
EUR 120,000 - 180,000
Senior AI Solutions Engineer
Senior AI Solutions Engineer

Morgan McKinley • Dublin

Hybrid
EUR 95,000 - 115,000
Competitive base salary
Pension and private healthcare
Hybrid work policy
AI Engineer – Generative AI / LLM / ML
AI Engineer – Generative AI / LLM / ML

The Recruitment Company Australia • Leinster

Hybrid
EUR 90,000 - 130,000
Senior AI Engineer | Senior AI Engineer
Senior AI Engineer | Senior AI Engineer

Johnson Controls Ireland HVAC • Cork

Hybrid
EUR 120,000 - 180,000
Artificial Intelligence Specialist
Artificial Intelligence Specialist

Archer Recruitment • Dublin

On-site
EUR 90,000 - 130,000
AI Architecture & Engineering Lead
AI Architecture & Engineering Lead

DB Recruitment • Dublin

On-site
EUR 120,000 - 180,000