Senior Data Scientist/AI Engineer (Reinforcement Learning)
HeadHR
Województwo mazowieckie
On-site
PLN 156,240 - 223,200
Full time
14 days+
Application generator
A complete application in a minute — tailored resume and cover letter, ready to send.
Get past ATS filters
Benefits offered by this job
Atrakcyjne wynagrodzenia
Możliwość pracy zdalnej
Udział w interesujących projektach
Job summary
HeadHR poszukuje doświadczonego inżyniera oprogramowania specjalizującego się w Pythonie oraz uczeniu maszynowym. Osoba na tym stanowisku będzie odpowiedzialna za projektowanie i wdrażanie środowisk dla agentów, tworzenie efektywnych rurociągów danych, a także współpracę z zespołem inżynierów w celu zapewnienia wydajności i skalowalności systemu. Oferujemy atrakcyjne wynagrodzenie, możliwość pracy zdalnej oraz udział w interesujących projektach.
Qualifications
Ponad 5-letnie doświadczenie w inżynierii oprogramowania w Pythonie.
Minimum 3 lata doświadczenia jako Data Scientist lub inżynier ds. uczenia maszynowego.
Praktyczna znajomość frameworków AI.
Responsibilities
Projektowanie i wdrażanie środowisk RL do oceny agentów.
Tworzenie rurociągów do generowania zadań i dynamicznych zbiorów danych.
Optymalizacja wydajności środowiska oraz logowanie nagród.
Skills
Doświadczenie w inżynierii oprogramowania w Pythonie
Znajomość frameworków AI (Langchain, Langraph, mcp-server)
Doświadczenie w pracy z sztuczną inteligencją
Job description
Requirements:
Over 5 years of experience in software engineering in Python.
At least 3 years of experience in the position of Data Scientist, Machine Learning/Environment Engineering.
Working hours from 2:00 PM to 10:00 PM.
Practical knowledge of AI frameworks (Langchain, Langraph, mcp-server).
Extensive practical experience in working with artificial intelligence, including instant engineering and climate coding.
Additional advantages:
Knowledge of the Code of Conduct or Claude's Code.
Experience in integrating artificial intelligence with the system will be an additional asset.
Understanding of RL concepts - reward modeling, environmental dynamics, verifiability, evaluation, and agent interaction loops.
Knowledge of tools, metrics, and data channels for evaluating RL.
Expertise in planning own work.
Kogo poszukujemy?
Responsibilities:
Designing and deploying RL environments for large-scale agent evaluation and reinforcement learning experiments.
Create pipelines for task generation, dynamic datasets, and scripted environments with controlled complexity and stochasticity.
Develop validators and reward models to automatically evaluate trajectories and assess model inference.
Collaborate with infrastructure and systems engineers to ensure scalability, reproducibility, and equip environments with tools for detailed telemetry.
Design API interfaces and orchestration structures for running, resetting, and evaluating agents in various environments.
Optimization of environment performance, reward logging, and reproducibility in distributed configurations.