Remote ML & NLP Expert for Evaluation & R&D

Weekday 1

United States

Remote

USD 110,000 - 152,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Weekday 1 is seeking experienced Machine Learning & NLP Experts to contribute to training and evaluating frontier AI models, designing challenging tasks, and creating high-quality reference solutions. You will assess model performance, identify gaps, and provide analyses in a remote, independent contractor role.

This part-time, fully remote opportunity requires about 20 hours per week and offers compensation of $80-$110 per hour, with weekly payments via Stripe or Wise.

Qualifications

  • Deep hands-on experience in Machine Learning and/or NLP through industry, research, or graduate/PhD-level work.
  • Strong proficiency in Python with practical experience developing ML or NLP applications.
  • Strong understanding of modern ML techniques including Model Training and Evaluation, Transformer Architectures, LLMs, NLP Pipelines, and Feature Engineering.
  • Experience with PyTorch, TensorFlow, Hugging Face Transformers, or equivalent.
  • Ability to commit approximately 20 hours per week.
  • Excellent written communication skills and the ability to work independently in a remote environment.

Responsibilities

  • Design challenging, real-world ML and NLP tasks covering ML model development, NLP, IR, and LLMs.
  • Develop accurate reference solutions and integrate tasks into agentic development environments using Python.
  • Build executable evaluation frameworks and testing components where appropriate.
  • Evaluate AI model outputs for technical correctness, reasoning quality, and overall performance.
  • Identify capability gaps, classify model failure modes, and provide detailed analyses.
  • Create and refine evaluation guidelines, scoring rubrics, and quality standards for ML and NLP tasks.
  • Collaborate with subject matter experts to ensure consistency, accuracy, and high-quality training data.

Skills

Machine Learning
Natural Language Processing
Python
PyTorch
TensorFlow
HuggingFace Transformers
LLMs
Model Evaluation
NLP Pipelines
RAG

Tools

PyTorch
TensorFlow
HuggingFace Transformers

Job description

Weekday 1 is seeking experienced Machine Learning & NLP Experts to contribute to training and evaluating frontier AI models, designing challenging tasks, and creating high-quality reference solutions. You will assess model performance, identify gaps, and provide analyses in a remote, independent contractor role.

This part-time, fully remote opportunity requires about 20 hours per week and offers compensation of $80-$110 per hour, with weekly payments via Stripe or Wise.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote ML Engineer - RLHF & Model Evaluation Expert
Remote ML Engineer - RLHF & Model Evaluation Expert

prolificacademicltd • United States

Remote
USD 55,000 - 110,000
Competitive hourly rates
Remote-first, flexible schedule
Lightweight onboarding
Machine Learning & NLP Expert
Machine Learning & NLP Expert

Weekday 1 • United States

Remote
USD 110,000 - 152,000
Remote ML Evaluation Architect (Part-Time)
Remote ML Evaluation Architect (Part-Time)

anyone-ai • United States

On-site
USD 83,000 - 124,000
Remote ML & AI Expert (PhD) for Model Evaluation
Remote ML & AI Expert (PhD) for Model Evaluation

Anyone AI Inc. • United States

Remote
USD 176,000 - 238,000
Senior AI/ML Engineer — Remote, Flexible Hours
Senior AI/ML Engineer — Remote, Flexible Hours

Hidden Jobs • United States

Remote
USD 66,000 - 110,000
Fully remote
Flexible hours
No minimum commitment
Senior AI/ML Evaluator (Remote, Per-Study)
Senior AI/ML Evaluator (Remote, Per-Study)

prolificacademicltd • United States

Remote
USD 83,000 - 110,000
Fully remote
Flexible hours
Short onboarding
Senior AI/ML Engineer - LLM Training & Evaluation (Remote)
Senior AI/ML Engineer - LLM Training & Evaluation (Remote)

prolificacademicltd • United States

Remote
USD 69,000 - 110,000
Hourly pay up to $80
Fully remote
Flexible hours
+1
Contract AI/ML Engineer — Remote RLHF Training & Evaluation
Contract AI/ML Engineer — Remote RLHF Training & Evaluation

prolificacademicltd • United States

Remote
USD 55,000 - 110,000
Flexible scheduling
Fully remote
Remote Senior AI/ML Auditor & Evaluator
Remote Senior AI/ML Auditor & Evaluator

prolificacademicltd • United States

Remote
USD 66,000 - 110,000
Remote work
Flexible scheduling
Senior AI Engineer - Remote Research & RLHF Evaluation
Senior AI Engineer - Remote Research & RLHF Evaluation

prolificacademicltd • United States

Remote
USD 66,000 - 110,000
Remote work
Flexible hours
Short onboarding