Senior ML Researcher

Toloka

España

A distancia

EUR 90.000 - 130.000

Jornada completa

14 días+
Generador de candidaturas

No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Remote work
Competitive compensation + ESOP
PTO & benefits

Descripción de la vacante

Toloka AI is a remote-first company powering GenAI models with a global team. You will own Toloka’s end-to-end fine-tuning, RL, and evaluation stack, turning complex ML experiments into scalable platform features and collaborating with engineers to deliver end-to-end solutions for clients.

The role spans data prep, model distillation, evaluation, and improving the platform through prompt/tool design and human-judged evaluation loops. English proficiency required.

Formación

  • 4+ years in ML engineering or applied research.
  • Hands-on with LLMs in production or research settings.
  • Experience fine-tuning open-weight models (LoRA/full FT) and data curation.
  • Strong Python and PyTorch; familiar with the Hugging Face ecosystem.

Responsabilidades

  • Own end-to-end fine-tuning pipelines: data prep, training, distillation, evaluation, serving.
  • Extend post-training stack into RL and reward modeling; design user-facing RL flow.
  • Build and calibrate evaluation harnesses with LLMs as judges against human labels.
  • Improve the platform's guiding agent: prompts, tools, eval loops.
  • Run experiments for client projects and translate results into platform features.
  • Collaborate with platform engineers on SDK/API surfaces for training/eval jobs.

Conocimientos

ML engineering
Applied research
LLMs in production
Python
PyTorch
HF ecosystem

Herramientas

vLLM
Transformers

Descripción del empleo

About Toloka

At Toloka AI we create data that powers leading GenAI models and innovations. We work with frontier labs, big tech, renowned AI startups, enterprises and non‑profit research organizations worldwide. We use a combination of Experts + Crowd + Tech Platform to teach AI models to reason and evaluate their efficacy and safety. We have experts in more than 50 different domains—from doctors and lawyers to physicists and engineers—and boast one of the most diverse global crowds,representing ove r 100 countries and speaking 40+ languages. We are a well‑funded startup with an enviable portfolio of clients including Anthropic, Amazon, Microsoft, Poolside, Recraft, and Shopify.

Recently, we secured strategic investment led by Bezos Expeditions and Nebius Group with participation from Mikhail Parakhin, CTO of Shopify and board advisor to leading GenAI companies, who now serves as our Chairman of the Board. Our remote‑first team is globally distributed around the world: USA, UK, the Netherlands, Serbia, and more.

About the Team

We are the ML team inside Toloka — we build the machine‑learning products that power the platform itself, so every project running on Toloka is faster, cheaper, and more reliable.

A few examples of what we own:

  • LLM QA — the core technology behind Toloka's automated quality‑check mechanism. Every annotation flowing through Self‑Service is reviewed by an LLM agent we design, train, and operate.
  • Model distillation and fine‑tuning — adapting frontier and open‑source models to Toloka's tasks to hit the right quality at the right cost.
  • Evaluation, benchmarking, cost modeling, and model selection across providers.

We own the full chain. The same team designs the ML solution, ships it to production, keeps it running 24/7, analyzes the results coming back from real projects, and feeds that signal into the next iteration. No hand‑off between research, engineering, and operations — it's all us.

About the Position

You will own Toloka’s end-to-end fine‑tuning, RL, and evaluation stack, bridging applied research and product engineering. In this role, you will spearhead greenfield post‑training initiatives (such as GRPO and reward modeling), transform complex ML experiments into scalable platform features, and occasionally author technical write‑ups on your findings for the AI community.

What you’ll do
  • Own end-to-end fine‑tuning pipelines: data prep, SFT/LoRA training, distillation from frontier models to smaller ones, evaluation, and serving handoff — both as self‑serve platform capabilities and in hands‑on client engagements.
  • Extend our post‑training stack beyond SFT into RL (RFT/GRPO-style methods, reward modeling, LLM‑judge‑based rewards) and help design the user‑facing RL flow on the platform — this part is greenfield.
  • Build and calibrate evaluation harnesses: LLM‑as‑judge setups calibrated against human labels, golden datasets, regression evals for optimization runs.
  • Improve the platform's guiding agent: prompt and tool design, eval‑driven improvement loops, stress‑testing scenarios and fixing what breaks.
  • Run experiments for client projects (e.g. prompt compression vs. fine‑tuning trade‑off studies) and turn the results into repeatable platform features.
  • Work closely with platform engineers on the SDK/API surface so that training and eval jobs are callable from a developer's existing workflow.
What we're looking for
  • 4+ years in ML engineering or applied research, with at least 1–2 years hands‑on with LLMs in production or research settings.
  • Practical experience fine‑tuning open‑weight models (LoRA/full FT), including data curation and knowing when fine‑tuning is the wrong answer.
  • Solid grasp of LLM evaluation: building evals from scratch, LLM‑as‑judge pitfalls, calibration against human judgments.
  • Strong Python engineering: you write code others can run, not just notebooks; comfortable with the training/inference stack (PyTorch, HF ecosystem, vLLM or similar).
  • Product mindset: you'll often be the ML person closest to a client problem, so you need to reason about what's worth building, not just what's possible.
  • Comfortable with ambiguity — priorities shift as we learn from pilots.
  • Language: Fluency in English (B2 or above).
Nice to have
  • Hands‑on RL for LLMs: GRPO/PPO‑style post‑training, reward modeling, RLHF/RLAIF pipelines.
  • Prompt optimization frameworks (DSPy/GEPA or similar) or prompt‑compression research (gisting).
  • Experience with distillation and quantization for cost/latency optimization.
  • Experience building agentic systems (tool use, multi‑step workflows) or shipping ML features in a self‑serve product.
What we can offer
  • You will be part of an international, dynamic environment that drives innovation and sets new standards in the AI and technology sector.
  • Competitive compensation package including base salary, bonus, and ESOP.
  • Paid PTO and benefits will vary depending on location.
  • We offer a full remote or hybrid model (if you are based in NL or Serbia).
  • IT setup and home office allowances.
Equal Opportunity Employer:

Toloka is committed to providing equal opportunity and fostering an inclusive environment. We welcome applications from all qualified individuals and do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, gender identity, age, marital status, veteran status, disability, or any other characteristic protected by applicable law. Selection decisions are based on qualifications, merit, and business need.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Program Director — Enterprise Fine-Tuning
Program Director — Enterprise Fine-Tuning

Toloka Ai • España

A distancia
EUR 90.000 - 140.000
Competitive compensation with base,–,+
ESOP and bonuses
Remote or hybrid work model
+1
Freelance Technical Solutions Engineer
Freelance Technical Solutions Engineer

Toloka Ai • España

Presencial
EUR 45.000 - 70.000
Remote work possible
Hybrid work option
Flexible scheduling
Senior ML Researcher: End-to-End Fine-Tuning & RL Innovation
Senior ML Researcher: End-to-End Fine-Tuning & RL Innovation

Toloka • España

A distancia
EUR 90.000 - 130.000
Remote work
Competitive compensation + ESOP
PTO & benefits
Freelance Data Partnerships Lead
Freelance Data Partnerships Lead

Toloka • España

A distancia
EUR 90.000 - 130.000
Fully remote
Flexible schedule
Global collaboration
+1
Mid/Senior AI Engineer
Mid/Senior AI Engineer

TensorOps Consulting Services Ltd • España

A distancia
EUR 60.000 - 90.000
Remote work
Certifications funded
Dynamic, high-impact projects
+4
Senior AI Engineer — Agentic ML & Autonomous Agents (Remote)
Senior AI Engineer — Agentic ML & Autonomous Agents (Remote)

Toloka • España

Presencial
EUR 70.000 - 110.000
Competitive compensation
Bonus
ESOP
+3
Senior MLOps Engineer (Training & Inference Optimization)
Senior MLOps Engineer (Training & Inference Optimization)

multiversecomputing • Donostia/San Sebastián

Presencial
EUR 70.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8
Backend Team Lead
Backend Team Lead

Acclaim AI • Barcelona

Presencial
EUR 90.000 - 130.000
Fully remote across Europe
Private English lessons via Preply
Company-paid subscriptions to top AI‑m
Ultralytics LLM Engineer
Ultralytics LLM Engineer

Tmoose • Madrid

Presencial
EUR 85.000 - 125.000
Competitive salary
24 days paid vacation
Home setup allowance
+2
Founding ML Engineer (Spectrum)
Founding ML Engineer (Spectrum)

JetBrains • Madrid

Presencial
EUR 70.000 - 90.000
Competitive salary
Generous runway and corporate resources