ML Research Engineer - PhD

Obsidian

Berlin

Vor Ort

EUR 60.000 - 90.000

Teilzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Obsidian is hiring experienced machine learning engineers and researchers to participate as baseliners for evaluations of AI agents on machine learning tasks. Candidates will work independently on ML research tasks under fixed conditions using their preferred tools.

The ideal candidates should have significant experience in machine learning, preferably from top-tier universities or companies like FAANG. The role requires a minimum commitment of 20 hours per week, with a preference for more availability.

Qualifikationen

  • 3+ years of machine learning experience required.
  • Attended a top-100 university or equivalent experience.
  • Deep expertise in ML frameworks essential.

Aufgaben

  • Attempt ML research tasks under fixed constraints.
  • Work independently in a sandboxed environment.
  • Submit final work products along with documentation.

Kenntnisse

Machine learning experience
Expertise in ML frameworks
Collaboration in evaluations

Ausbildung

Education from top-100 university or FAANG experience

Tools

PyTorch
JAX
TensorFlow

Jobbeschreibung

Overview

We are hiring experienced machine learning engineers and researchers to serve as human baseliners for evaluations of open-ended machine learning research tasks. These evaluations measure how well AI agents perform on realistic AI R&D problems. To interpret agent performance, we also need strong human reference points: skilled practitioners attempting the same tasks under the same time and compute constraints. As a baseliner, you will complete self-contained ML research tasks in a sandboxed environment, working independently with your preferred tools and workflow. Your performance will be used as a benchmark against which frontier-model agents are evaluated.

What You’ll Do
  • Attempt open-ended machine learning research tasks under a fixed time and compute budget (work trial)
  • Work independently in a sandboxed Linux environment with internet access
  • Use your preferred tooling, including IDEs and AI coding assistants such as Cursor, Claude Code, and ChatGPT
  • Record your full working session via screen recording
  • Complete a short pre-task and post-task questionnaire
  • Submit your final work product, screen recording, and completed questionnaires

Post this you will be hired for a longer commitment.

Commitment
  • Minimum 20 hours per week if selected
  • More availability is strongly preferred
Requirements
  • 3+ years of machine learning experience (time spent in a PhD program counts toward this requirement; undergraduate and master’s experience does not count)
  • Attended a top‑100 university or worked at FAANG or a comparable company
  • Experience with at least one major ML framework such as PyTorch, JAX, or TensorFlow
  • Deep, hands‑on expertise in at least one of the following focus areas:
  • Pretraining under tight data and compute budgets
  • PPO, reward shaping, custom gym / gymnasium environments, and throughput tuning
  • Full fine‑tuning, LoRA, QLoRA, DPO, RLHF, RLAIF, and distillation
  • Large‑scale corpus filtering, deduplication, subsampling, and benchmark contamination avoidance
  • Architecture design under strict parameter‑count or size constraints
  • Modifying pretrained architectures, including attention patterns, pooling heads, or training objectives
  • Contrastive training for embedding or retrieval models
  • Generative vision or video modeling
  • Multilingual or low‑resource language experience
  • Image or video data pipelines at scale
  • Experience balancing competing model objectives such as safety and capability
  • Prior work as an ML evaluator, red‑teamer, or baseliner
Required Domain Expertise
  • Pretraining: training transformer language models from scratch
  • Reinforcement learning: training agents in custom or existing environments
  • Post‑training: fine‑tuning and aligning LLMs
  • Dataset curation: building and cleaning large text corpora for LLM training
  • Model architecture: designing and modifying neural network architectures
Logistics (work trial requirements)
  • One baseline attempt per contractor per task
  • Each task may only be attempted once by a given contractor
  • All work is confidential and covered by NDA
  • Compute and environment are provided; no personal GPU is required
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Research Engineer (Agentic Models)
Research Engineer (Agentic Models)

United States Digital Space LLC • Berlin, München

Hybrid
EUR 90.000 - 120.000
Research Engineer (Agentic Models)
Research Engineer (Agentic Models)

JetBrains • München

Vor Ort
EUR 70.000 - 95.000
ML Engineer - Coding Agent Expert
ML Engineer - Coding Agent Expert

Obsidian • Berlin

Vor Ort
ML Engineer (Coding Agent Experience)
ML Engineer (Coding Agent Experience)

aitrainer • Deutschland

Remote
ML Engineer - AI Coding Expert
ML Engineer - AI Coding Expert

Obsidian • Berlin

Vor Ort
EUR 394.000 - 561.000
AI/ML Engineer - Fully Remote | Upto $85/hr
AI/ML Engineer - Fully Remote | Upto $85/hr

Obsidian • Berlin

Remote
EUR 26.000 - 45.000
ML Scientist - Adversarial Robustness
ML Scientist - Adversarial Robustness

Obsidian • Berlin

Vor Ort
EUR 90.000 - 140.000
ML Scientist - Adversarial Robustness
ML Scientist - Adversarial Robustness

Mercor • Berlin

Vor Ort
EUR 90.000 - 150.000
Flexible project-based work
Competitive compensation
Staff Research Engineer (LLM Pre-Training)
Staff Research Engineer (LLM Pre-Training)

United States Digital Space LLC • Berlin, München

Hybrid
EUR 90.000 - 140.000
Research Engineer (Post-Training / RL) Munich or Zurich
Research Engineer (Post-Training / RL) Munich or Zurich

MicroAGI, Inc. • München

Vor Ort
EUR 70.000 - 100.000
MuJoCo
Isaac Sim
ManiSkill