Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
nineDots.io in Dublin is seeking an AI Research Engineer focused on frontier models and benchmarking. You will design and run rigorous evaluation protocols across multiple state-of-the-art LLMs and analyze model behavior.
The role values deep research credentials, publications or demonstrated benchmark work, and the ability to build evaluation systems and datasets. You’ll work on post-training and reinforcement learning experiments and compare model strengths across several families.
AI Research Engineer – Frontier Models & Benchmarking
Dublin | 5 Days Onsite
We’re working with a well-funded AI company building training and evaluation environments for some of the world’s leading AI labs.
They’re growing a small research team in Dublin and are looking for people who are genuinely deep into frontier AI models.
This is a deliberately specialised role.
If your LLM experience is mainly building RAG applications, chatbots, agent orchestration or integrating models into existing products, this probably isn’t the right role for you.
They’re looking for people who study the models themselves.
You’ll be working on problems like:
The people we particularly want to hear from have experience in one or more of:
Strong research credentials are highly valued. That could mean a PhD in a relevant area, significant research experience, publications at conferences such as NeurIPS, ICML or ICLR, or demonstrably strong work building benchmarks and evaluation systems.
Most importantly, we’re looking for people who have actually done this work, rather than simply used the terminology.
If your experience is predominantly:
...without deeper model evaluation, benchmarking or research experience, this role is unlikely to be a fit.
This is also 5 days per week onsite in Dublin city centre. It’s a small, ambitious team that moves quickly and works hard, so it’ll suit someone who actively wants that kind of start-up environment rather than a traditional 9-to-5.
If you’re the sort of person who sees a new frontier model released and immediately wants to test it, break it, compare it and understand why it behaves differently from the others, we’d like to hear from you.