Machine Learning Expert - Fully Remote | Upto $90/hr

Obsidian

San Francisco (CA)

Hybrid

USD 120,000 - 160,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Obsidian in San Francisco is hiring experienced machine learning engineers and researchers to evaluate AI performance on real-world tasks. Candidates will work independently in a sandboxed environment while completing ML research tasks under time constraints.

A minimum of 20 hours per week is expected, with more availability preferred. Applicants must have 3+ years of machine learning experience and ideally have attended a top-100 university or worked at FAANG.

Qualifications

  • 3+ years of machine learning experience.
  • Strong expertise in ML frameworks like PyTorch or TensorFlow.
  • Worked at a top-100 university or FAANG company.

Responsibilities

  • Attempt open-ended ML research tasks under fixed budget.
  • Work independently in a sandboxed environment.
  • Record sessions and submit work product.

Skills

Machine Learning Experience
Experience with ML Frameworks (PyTorch, JAX, TensorFlow)
Data Pretraining
Reinforcement Learning
Model Architecture Design

Education

Degree from a top-100 university or FAANG equivalent

Job description

Overview

We are hiring experienced machine learning engineers and researchers to serve as human baseliners for evaluations of open-ended machine learning research tasks. These evaluations measure how well AI agents perform on realistic AI R&D problems. To interpret agent performance, we also need strong human reference points: skilled practitioners attempting the same tasks under the same time and compute constraints. As a baseliner, you will complete self-contained ML research tasks in a sandboxed environment, working independently with your preferred tools and workflow. Your performance will be used as a benchmark against which frontier-model agents are evaluated.

What You’ll Do
  • Attempt open-ended machine learning research tasks under a fixed time and compute budget (work trial)
  • Work independently in a sandboxed Linux environment with internet access
  • Use your preferred tooling, including IDEs and AI coding assistants such as Cursor, Claude Code, and ChatGPT
  • Record your full working session via screen recording
  • Complete a short pre-task and post-task questionnaire
  • Submit your final work product, screen recording, and completed questionnaires

Post this you will be hired for a longer commitment.

Commitment
  • Minimum 20 hours per week if selected
  • More availability is strongly preferred
Requirements
  • 3+ years of machine learning experience (time spent in a PhD program counts toward this requirement; undergraduate and master’s experience does not count)
  • Attended a top‑100 university or worked at FAANG or a comparable company
  • Experience with at least one major ML framework such as PyTorch, JAX, or TensorFlow
  • Deep, hands‑on expertise in at least one of the following focus areas:
  • Pretraining under tight data and compute budgets
  • PPO, reward shaping, custom gym / gymnasium environments, and throughput tuning
  • Full fine‑tuning, LoRA, QLoRA, DPO, RLHF, RLAIF, and distillation
  • Large‑scale corpus filtering, deduplication, subsampling, and benchmark contamination avoidance
  • Architecture design under strict parameter‑count or size constraints
  • Modifying pretrained architectures, including attention patterns, pooling heads, or training objectives
  • Contrastive training for embedding or retrieval models
  • Generative vision or video modeling
  • Multilingual or low‑resource language experience
  • Image or video data pipelines at scale
  • Experience balancing competing model objectives such as safety and capability
  • Prior work as an ML evaluator, red‑teamer, or baseliner
Required Domain Expertise
  • Pretraining: training transformer language models from scratch
  • Reinforcement learning: training agents in custom or existing environments
  • Post‑training: fine‑tuning and aligning LLMs
  • Dataset curation: building and cleaning large text corpora for LLM training
  • Model architecture: designing and modifying neural network architectures
Logistics (work trial requirements)
  • One baseline attempt per contractor per task
  • Each task may only be attempted once by a given contractor
  • All work is confidential and covered by NDA
  • Compute and environment are provided; no personal GPU is required
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Research Engineer - PhD - AI Trainer
ML Research Engineer - PhD - AI Trainer

Obsidian • Seattle (WA)

Hybrid
USD 100,000 - 150,000
ML Engineer Specialist - Freelance AI Trainer Project
ML Engineer Specialist - Freelance AI Trainer Project

Meridial • United States

On-site
Flexibility to set your schedule
Remote work environment
Impact on cutting-edge AI
Applied AI Engineer
Applied AI Engineer

SherlockTalent • Miami (FL)

Hybrid
USD 120,000 - 140,000
Solid Benefits
Referral bonus of $2,500
Senior Staff Engineer (Machine Learning) - 45391
Senior Staff Engineer (Machine Learning) - 45391

Turing • Town of Italy (NY)

Remote
USD 120,000 - 160,000
Fully remote environment
Opportunity to work on cutting-edge AI projects
Flexible working hours
Senior Staff Engineer (Machine Learning) - 45391
Senior Staff Engineer (Machine Learning) - 45391

Turing • Germany (OH)

Remote
USD 110,000 - 165,000
Work in a fully remote environment
Opportunity to work on cutting-edge AI projects
4h/day min
+4
Machine Learning Developer (Freelance)
Machine Learning Developer (Freelance)

Mindrift • Wisconsin

On-site
USD 102,000 - 146,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required
Machine Learning Developer (Freelance)
Machine Learning Developer (Freelance)

Mindrift • Missouri

On-site
USD 74,000 - 124,000
Machine Learning Research Engineer
Machine Learning Research Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
ML Engineer - Coding Agent Expert
ML Engineer - Coding Agent Expert

Obsidian • New York (NY)

Hybrid