Machine Learning Lead (AI Data Labeling)

NewtonX

United States

On-site

USD 140,000 - 190,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401K match
Paid vacation and holidays
Paid sick days
Parental leave
Office snacks

Job summary

NewtonX is expanding into AI data annotation and RLHF. This role is the technical owner of data quality; you will partner with the Program Lead, translating client goals into concrete technical specs that guide data design and evaluation.

You will be hands-on, shaping task environments and rubric structures for advanced AI workflows. You will validate datasets at the ML level, diagnose data quality issues, and defend technical choices to clients and the ops team.

Qualifications

  • 2+ years hands-on RL experience
  • Strong Python and model API familiarity
  • Ability to translate customer requirements into technical specs

Responsibilities

  • Design the task, environment, and rubric structure for agentic workflows
  • Own ML-level validation of datasets before delivery
  • Assess data quality across training and evaluation pipelines
  • Defend data quality decisions to clients and internal teams

Skills

Python
RLHF
SFT
RLVR
Agentic RL
CoT
Model APIs
Data quality
Statistical rigor

Tools

Eval harness

Job description


  • NewtonX is rapidly expanding into the AI data annotation and RLHF (Reinforcement Learning from Human Feedback) space

  • We are leveraging our core superpower—recruiting the world’s leading domain experts—to provide high-quality, expert-led data labeling for AI labs and enterprises

  • Beyond our unmatched B2B recruiting, we utilize powerful, automated project management processes that allow us to scale projects rapidly, adapt to shifting requirements, and manage our subject matter experts with professional, industry-best practices and fair compensation

  • This role is the technical owner of data quality

  • You will partner directly with the Program Lead and act as technical lead in communicating with clients, interpreting their core AI model testing goals and assisting the Program Lead in creating concrete technical specs that will accomplish these goals

  • The core question you own: will the data we produce do its job once the customer trains on it or evaluates with it? Clean data that passes every operational gate can still fail — a training set that yields a weak or misleading signal, or an eval that technically runs but doesn’t surface the weaknesses that matter. You are the person who can look at an approved submission and say “this is technically correct and still won’t do the job, and here’s why.”

  • You are hands-on and close to the work. This is a foundational role — what it covers will grow as the business does

  • Own the judgment of whether task designs and rubrics produce a useful training signal for the consuming method (SFT, RLHF, RLVR, agentic RL, CoT, evals). Catch mismatches between what the data rewards and what the customer is actually training for

  • Design the task, environment, and rubric structure for agentic workflows

  • Own the ML-level validation of the dataset before delivery — not just whether individual submissions meet spec

  • Partner with the Program Lead to convert customer requirements into concrete technical specs: expert profiles, screener trees, task interfaces, task templates, QC rubrics, statistical thresholds

  • Partner with the Program Lead to define and defend quality metrics: inter-annotator agreement targets, gold-standard injection rates, statistical power thresholds — providing the statistical and methodological grounding

  • Calibrate the ops team on what “good” looks like per engagement; run alignment sessions when standards shift

  • Serve as the technical counterpart to the customer’s ML, applied science, and product teams. Hold your ground on the technical questions that decide data quality across training and evaluation — task and reward design, data quality for SFT and preference methods, RL and RLVR, contamination, statistical rigor, and agentic workflows

  • Help diagnose data-quality questions when a customer reports the data underperformed — reason through whether the issue is the data, the quantity, or something in their training setup, and make the case defensibly


Benefits


  • Medical, dental, and vision insurance

  • 401K 3% match, immediate vesting

  • Paid Vacation and Public Holidays

  • Paid Sick Days

  • Pre-tax commuter benefits

  • Health Savings / Flexible Savings Account

  • Paid Parental / Family Leave

  • Office snacks and refreshments

  • Lunch and learns

  • Monthly team outings and bimonthly happy hours

  • Annual company retreat

  • Volunteering

  • Virtual fun, social activities (e.g., happy hours, painting, escape rooms, trivia, cooking classes, meditation, and more!)


If the profile above describes you and your passions, we’d love to hear from you!



  • Strong programming foundation: read and reason about an eval harness, write Python comfortably, work with model APIs, prototype scoring pipelines. Not a production engineer, but not hands-off

  • At least 2 years working hands-on with RL, including how tool-use trajectories are rewarded and evaluated. If you’re not fluent in RL, this isn’t the role — it’s the foundation of the core judgment you’ll be making

  • Strong written communication: methodology sections, technical reports, and specs that hold up to expert review

  • Working fluency across modern LLM post-training and evaluation: SFT, RLHF/preference data quality, RLVR, chain-of-thought, eval harness construction, contamination handling, statistical significance, and agentic/tool-use evaluation

  • Statistical fluency: you know when an effect is real vs. noise and can defend a sample size or significance threshold

  • Client-facing presence: you’ve defended technical design choices in real time to skeptical audiences and adjusted scope without losing rigor. Range matters — you can talk to a Series B CTO and a Fortune 100 AI lead in the same week

  • Deep applied ML experience centered on post-training human data — you’ve owned a human-data workstream as an applied scientist or ML engineer

  • Genuine understanding of how training data becomes model behavior — you can reason about what a model will learn from a given dataset, not just whether the data meets spec

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Engineer
Applied AI Engineer

SherlockTalent • Miami (FL)

Hybrid
USD 120,000 - 140,000
Solid Benefits
Referral bonus of $2,500
AI Research Scientist, Learning & Evaluation
AI Research Scientist, Learning & Evaluation

Studyfetch • Beverly Hills (CA)

On-site
USD 150,000 - 210,000
Medical, Dental, Vision (100% employer
75% dependent coverage
401(k) with employer matching
+2
ML Lead, AI Data Labeling
ML Lead, AI Data Labeling

NewtonX • United States

Remote
USD 120,000 - 160,000
Comprehensive Benefits
401k match with immediate vesting
Paid time off
AI Research Scientist, Learning & Evaluation
AI Research Scientist, Learning & Evaluation

Socket.dev • Beverly Hills (CA)

On-site
USD 180,000 - 280,000
Daily team dinner provided in-office
Member of Technical Staff [AI/ML Engineer]
Member of Technical Staff [AI/ML Engineer]

Burnt Group • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 275,000
Equity
Member of Technical Staff, Post-Training
Member of Technical Staff, Post-Training

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Meals and office benefits
Visa sponsorship
Senior Engineering Manager
Senior Engineering Manager

Mixpeek • San Francisco (CA)

On-site
USD 170,000 - 250,000
Training Specialist
Training Specialist

Enterprise • San Francisco (CA)

On-site
USD 75,000 - 95,000
Research Engineer
Research Engineer

Cerebras • San Jose (CA)

On-site
USD 180,000 - 240,000
Applied Scientist
Applied Scientist

Vecna AI • Chicago (IL)

On-site
USD 140,000 - 210,000