Remote AI Research Scientist — Applied LLM Evaluation

AIToolboard

New York (NY)

Remote

USD 41,328 - 68,880

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Rex.zone is seeking an AI Research Scientist to lead applied AI research projects for US-based customers, translating open-ended questions into measurable experiments in LLM evaluation and RLHF data design. You will evaluate prompts, design datasets, and work with cross-functional teams to improve model performance, safety, and usefulness.

The role is remote in the United States, with compensation based on hourly pay and flexible collaboration across distributed teams.

Qualifications

  • Mid-Senior level with applied ML research or production ML evaluation.
  • Strong Python skills and PyTorch experience.
  • Hands-on LLM evaluation, prompt evaluation, or RLHF.
  • Experience in experiment design and metrics interpretation.
  • Familiar with dataset development, labeling, QA evaluation, and guidelines.
  • Excellent written communication for research artifacts.

Responsibilities

  • Own end-to-end applied research cycles: problem framing, baselines, ablations, and reporting.
  • Build and evaluate LLM systems using offline metrics and human-in-the-loop evaluation.
  • Design RLHF workflows: preference data specs, rater instructions, prompt sets, and rubric-based grading.
  • Create evaluation datasets and test suites: prompt evaluation, red-teaming prompts, and content safety labeling protocols.
  • Collaborate with data labeling teams on taxonomy, edge-case coverage, and training data quality.
  • Perform error analysis and model debugging to improve robustness, safety, and helpfulness.
  • Document methodology and results for reproducibility and auditability.

Skills

Python
PyTorch
LLM evaluation
RLHF
Experiment design
Data labeling
Communication

Job description

Rex.zone is seeking an AI Research Scientist to lead applied AI research projects for US-based customers, translating open-ended questions into measurable experiments in LLM evaluation and RLHF data design. You will evaluate prompts, design datasets, and work with cross-functional teams to improve model performance, safety, and usefulness.

The role is remote in the United States, with compensation based on hourly pay and flexible collaboration across distributed teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Research Scientist (United States, Remote)
AI Research Scientist (United States, Remote)

Rex.zone • United States

Remote
Remote AI Data Annotation & Evaluation Specialist
Remote AI Data Annotation & Evaluation Specialist

Rex.zone • Detroit (MI)

On-site
Remote AI Engineer (United States)
Remote AI Engineer (United States)

Rex.zone • United States

On-site
USD 100,000 - 140,000
Remote Applied AI Research Scientist (LLM & Evaluation)
Remote Applied AI Research Scientist (LLM & Evaluation)

Rex.zone • United States

Remote
USD 80,000 - 100,000
STEM Careers in the United States
STEM Careers in the United States

Rex.zone • United States

Remote
Remote AI Data Labeling & Evaluation Specialist
Remote AI Data Labeling & Evaluation Specialist

Rex.zone • United States

On-site
USD 42,000 - 68,000
Remote AI Research Scientist: LLM Evaluation & Experiments
Remote AI Research Scientist: LLM Evaluation & Experiments

Rex.zone • United States

Remote
USD 60,000 - 80,000
Senior Data Annotator - Remote AI QA & LLM Evaluation
Senior Data Annotator - Remote AI QA & LLM Evaluation

Rex.zone • Miami (FL)

On-site
Remote AI Engineer: RLHF & Evaluation Pipelines
Remote AI Engineer: RLHF & Evaluation Pipelines

Rex.zone • United States

On-site
USD 100,000 - 140,000
Remote AI/ML Research Engineer — LLM Training & Evaluation
Remote AI/ML Research Engineer — LLM Training & Evaluation

Rex.zone • United States

Remote
USD 80,000 - 100,000