Remote AI Research & Evaluation Lead

24-MAG

United States

Remote

USD 600,000 - 2,000,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Fully remote

Job summary

24-MAG LLC offers a full-time, fully remote opportunity at the intersection of AI research and real-world model performance. This role leads research evaluation, ML-oriented data design, failure analysis, and quality calibration to drive defensible signals and measurable improvements in AI systems.

The ideal candidate combines strong judgment on signal quality with experience in constructing evaluation frameworks and QA processes, while communicating findings clearly to stakeholders across

Qualifications

  • Strong judgment on research signal quality and readiness to support conclusions.
  • Experience designing ML-oriented datasets and evaluation frameworks.
  • Ability to translate complex real-world behavior into structured research opportunities.
  • Comfort making decisions in uncertain, evolving environments.
  • Excellent written and verbal communication skills.

Responsibilities

  • Own research and evaluation initiatives from framing to data design and signal validation.
  • Define rigorous approaches to ensure reliable, defensible research signal.
  • Analyze model and system failures to identify root causes and improvements.
  • Evaluate datasets, experiments, and conclusions against quality thresholds.
  • Collaborate with researchers, domain experts, and operators to align efforts.

Skills

Research signal quality
ML evaluation frameworks
Data design
Quality assurance
Communication skills
Systems-level thinking

Job description

24-MAG LLC offers a full-time, fully remote opportunity at the intersection of AI research and real-world model performance. This role leads research evaluation, ML-oriented data design, failure analysis, and quality calibration to drive defensible signals and measurable improvements in AI systems.

The ideal candidate combines strong judgment on signal quality with experience in constructing evaluation frameworks and QA processes, while communicating findings clearly to stakeholders across

Get your free, confidential resume review.

or drag and drop your file here.