AI Engineer (RL Environments)

Hashlist

Helsinki

On-site

EUR 80,000 - 120,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Meaningful equity
Central office in Helsinki
Work with model providers

Job summary

Hashlist expands its AI research team in Helsinki to develop domain-specific RL environments for constrained embedded programming and enterprise workflows. You will build and automate platforms for RL environments and simulate worlds to reveal model weaknesses.

You will translate training objectives into concrete data and evaluation specs, construct the reward layer, and run large-scale sandbox rollouts while collaborating with OEMs and AI Labs on fine-tuning challenges.

Qualifications

  • Production ML depth with Python and PyTorch for RL training.
  • End-to-end RL environment built and model trained.
  • Experience with containerization tools like Docker and scaling systems.
  • Strong communication and motivation.

Responsibilities

  • Build and automate our platform for creating RL environments.
  • Construct simulated worlds and expose model failure modes.
  • Turn AI objectives into concrete data and evaluation specs.
  • Build the reward layer and mitigation for reward hacking.
  • Run rollouts at scale with hundreds of sandboxed attempts.
  • Fine-tune open-source models and export datasets.
  • Communicate with OEMs and AI Labs on research questions.

Skills

Python
PyTorch
Reinforcement learning
TRL
OpenRLHF
LoRA fine-tuning
vLLM
SGLang
PPO
Docker

Tools

Docker

Job description

Would you like to operate at the frontier of AI evaluation, post-training, and model improvements?

We are now expanding our core Hashlist AI research team to support the creation of domain-specific RL environments for constraint-based embedded programming & complex enterprise engineering workflows.

What you will do:
  • Build and automate our platform for creating RL environments
  • Construct simulated worlds and explore data shapes that expose meaningful model failure modes across the embedded coding domain & related enterprise workflows
  • Turn AI training objectives into concrete data and evaluation specifications
  • Build the reward layer + reward-hacking mitigation
  • Run rollouts at scale. Hundreds of sandboxed attempts per task in parallel
  • Fine-tune open-source models: Before-and-after fine-tunes, failure reports, and dataset exports
  • Communication with our clients (OEMs & AI Labs) on specific research or fine-tuning questions they have
Skills needed:
  • Production machine learning depth. Python and PyTorch, reinforcement learning training with TRL, verl, SkyRL or OpenRLHF, PPO, GRPO or DPO, LoRA fine-tuning, and inference with vLLM or SGLang.
  • Training or evaluation systems, including at least one RL environment you built end-to-end and trained a model against. Comfortable with the infrastructure around it, e.g Docker, or similar containerization tools, to design and monitor systems at scale
  • Experience with harness & agentic optimisation for evals
  • Strong familiarity with common reinforcement learning algorithms and methods, especially with respect to post-training LLMs
  • High level of personal drive, motivation, and good communication skills.
Bonus:
  • You have published an environment, benchmark or evaluation harness we can look at.
  • You have worked with embedded, safety-critical or other physical engineering software.
Company benefits
  • Competitive compensation + meaningful equity
  • Central office in Helsinki
  • Be a part of a quickly scaling tech company working directly with model providers
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Embedded Platform Engineer
Embedded Platform Engineer

Hashlist • Helsinki

On-site
EUR 140,000 - 200,000
Competitive compensation and equity
Central office in Helsinki
Equity and growth opportunities
AI Engineer, RL Environments — Scale Sandboxed Training (Equity)
AI Engineer, RL Environments — Scale Sandboxed Training (Equity)

Hashlist • Helsinki

On-site
EUR 80,000 - 120,000
Competitive compensation
Meaningful equity
Central office in Helsinki
+1
Embedded Platform Engineer: RL Environments for AI
Embedded Platform Engineer: RL Environments for AI

Hashlist • Helsinki

On-site
EUR 140,000 - 200,000
Competitive compensation and equity
Central office in Helsinki
Equity and growth opportunities
Senior AI Engineer Applied AI · Loimaa / Helsinki, Finland · Full-time €75,000 – €110,000 + equ[...]
Senior AI Engineer Applied AI · Loimaa / Helsinki, Finland · Full-time €75,000 – €110,000 + equ[...]

Orex Nova, Inc. • Helsinki

On-site
EUR 90,000 - 130,000
Six weeks vacation
Annual learning budget
Equity in employee-owned company
AI Engineer
AI Engineer

3stepIT • Helsinki

On-site
EUR 65,000 - 90,000
Recognition & Rewards
Continuous Learning
Wellbeing Focus
+2
Senior AI Engineer Applied AI · Loimaa / Helsinki, Finland · Full-time €75,000 – €110,000 + equ[...]
Senior AI Engineer Applied AI · Loimaa / Helsinki, Finland · Full-time €75,000 – €110,000 + equ[...]

Orexnova • Helsinki

On-site
EUR 90,000 - 140,000
Six weeks vacation
€4,000 learning budget annually
Equity in an employee-owned company
Senior AI Engineer
Senior AI Engineer

Futurice Oy • Helsinki

On-site
EUR 90,000 - 120,000
Learning opportunities
Project variety
AI Developer - LLM Applications
AI Developer - LLM Applications

Reaktor • Helsingin seutukunta

Hybrid
EUR 90,000 - 130,000
Flexible work approach
Strong community support
Sustainable work-life balance
+2
ML/AI Engineer
ML/AI Engineer

Renessai • Helsinki

On-site
EUR 60,000 - 90,000
AI Scientist
AI Scientist

Brillian • Helsinki

Hybrid
Equity opportunities
Flexible working environment