Remote Research Engineer - AI Safety & Frontier Models

Singapore AI Safety Hub

Poland

Remote

PLN 448,000 - 673,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive benefits
Leave policies

Job summary

Neo Research (新衡) is seeking Research Engineers to build experimental systems for evaluating frontier models. You will create realistic agent environments, integrate tools, and develop long-horizon evaluation infrastructure, ensuring reliability and reproducibility across complex experiments.

You will collaborate with researchers, shaping evaluation methodology and turning open-ended questions into tractable experiments, with output directly contributing to published safety research.

Qualifications

  • Strong Python and general software-engineering skills.
  • Experience working with language models, agentic systems, or model evaluations.
  • Ability to turn an underspecified research question into a reliable experimental setup.
  • Strong engineering practices, particularly around reproducibility, testing, observability, and data integrity.
  • Ability to investigate unexpected results across both the model and the surrounding infrastructure.
  • Clear technical writing and communication skills.
  • Comfort working in an early-stage environment where requirements and research directions evolve.

Responsibilities

  • Design, implement, and run evaluations of misalignment and loss-of-control risks in frontier models.
  • Build realistic agent environments, scaffolds, and tool integrations for studying behaviour under increasing autonomy.
  • Develop infrastructure for long-horizon experiments involving many model calls, actions, tools, and environment states.
  • Own experimental reliability and reproducibility across model access, sampling, configuration, environment management, logging, and analysis.
  • Work with research scientists to turn open-ended questions into tractable experiments, and identify cases where an apparent model behaviour is actually an artefact of the evaluation.
  • Build tools for inspecting and analysing large collections of model trajectories and comparing behaviour across models and experimental conditions.
  • Contribute experimental methodology, infrastructure, and technical analysis directly to published research.

Skills

Python
Software engineering
Language models
Experimentation
Observability
Testing
Data integrity
Technical writing

Tools

Docker
Kubernetes

Job description

Neo Research (新衡) is seeking Research Engineers to build experimental systems for evaluating frontier models. You will create realistic agent environments, integrate tools, and develop long-horizon evaluation infrastructure, ensuring reliability and reproducibility across complex experiments.

You will collaborate with researchers, shaping evaluation methodology and turning open-ended questions into tractable experiments, with output directly contributing to published safety research.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Frontier AI Safety Research Scientist
Remote Frontier AI Safety Research Scientist

Singapore AI Safety Hub • Poland

Remote
PLN 561,000 - 747,000
Competitive benefits
Paid leave
Remote Lead Frontier AI Safety Scientist
Remote Lead Frontier AI Safety Scientist

Singapore AI Safety Hub • Poland

Remote
PLN 747,000 - 934,000
Competitive benefits
Senior AI Evaluation Architect for Frontier Models
Senior AI Evaluation Architect for Frontier Models

NVIDIA Corporation • Poland

Remote
PLN 375,000 - 650,000
Senior Frontier AI Safety Red Teamer
Senior Frontier AI Safety Red Teamer

Obsidian • Warszawa

Remote
PLN 250,000 - 420,000
Frontier AI Safety Evaluator
Frontier AI Safety Evaluator

Obsidian • Warszawa

On-site
PLN 180,000 - 300,000
Senior AI Safety Evaluator for Frontier Models
Senior AI Safety Evaluator for Frontier Models

Mercor • Warszawa

Remote
PLN 180,000 - 250,000
Senior AI Evaluation Engineer for Frontier Models
Senior AI Evaluation Engineer for Frontier Models

NVIDIA • Warszawa

On-site
PLN 390,000 - 650,000
Frontier AI Safety Red Teamer
Frontier AI Safety Red Teamer

Mercor • Warszawa

Remote
PLN 180,000 - 240,000
Frontier AI Safety Red Team Lead
Frontier AI Safety Red Team Lead

Mercor • Warszawa

On-site
PLN 180,000 - 300,000
Senior AI Safety Evaluator
Senior AI Safety Evaluator

Obsidian • Warszawa

Remote
PLN 120,000 - 180,000