AI Evaluation Specialist - RLHF, SFT & MCP (Remote, Immediate Joiners)

OpsLabel

India

Remoto

INR 1.653.000 - 3.306.000

Tempo pieno

31 ore fa
Candidati tra i primi
Generatore di candidature

Distinguiti per questa posizione — genera un curriculum e una lettera di presentazione personalizzati in circa un minuto.

Supera i filtri ATS

Descrizione del lavoro

OpsLabel is expanding its expert network and seeks AI evaluation specialists for RLHF, SFT, and MCP evaluation. You will judge AI outputs, stress test reasoning, and tool use against detailed rubrics on frontier models.

Role is remote with India preference, freelance/contract for long-term projects. Strong technical minds in ML, software, linguistics, or STEM welcome; advanced degree a plus but not required.

Competenze

  • Strong judgment on the quality, correctness, and safety of AI model outputs.
  • Comfort following detailed rubrics and guidelines.
  • Technical or analytical background in ML, software, data, linguistics, or STEM.

Mansioni

  • Rank and compare model responses to build high-quality preference data.
  • Write and review instruction-response and chain-of-thought examples.
  • Evaluate agent tool-use: tool-calling accuracy and multi-step workflows.
  • Catch subtle errors and reasoning failures.
  • Deliver high-quality work across multiple sprints with feedback from senior reviewers.

Conoscenze

Strong judgment
Technical/analytical background
Prompt engineering

Strumenti

RLHF
SFT
MCP
Tool calling

Descrizione del lavoro

AI Evaluation Specialist - RLHF, SFT & MCP (Remote, Immediate Joiners)

OpsLabel is expanding its expert network, and we have an immediate need for AI evaluation specialists. We need your judgment on frontier AI projects spanning RLHF, SFT, and MCP (agent tool-use) evaluation, helping train and stress-test how well models respond, reason, and use tools.

  • Remote (India preferred)
  • Freelance / Contract, long-term project work
  • Open to strong technical minds: ML, engineering, linguistics, or a STEM domain.
  • Advanced degree a plus, not required. What matters is sound judgment on what a good AI answer looks like.

Must have:

  • Strong judgment on the quality, correctness, and safety of AI model outputs
  • Comfort working to detailed rubrics and guidelines
  • A technical or analytical background (ML, software, data, linguistics, or a STEM field)

Nice to have:

  • Hands‑on experience with RLHF, preference ranking, or reward modeling
  • SFT data creation: instruction-response pairs or chain‑of‑thought
  • Familiarity with MCP, function or tool calling, or agent workflows
  • Programming or prompt‑engineering background

What you’ll do:

  • Rank and compare model responses to build high‑quality preference data
  • Write and review instruction‑response and chain‑of‑thought examples
  • Evaluate agent tool‑use: tool‑calling accuracy and multi‑step workflows
  • Catch the subtle errors and reasoning failures generic reviewers miss
  • Deliver high‑quality work across multiple sprints with direct feedback from senior reviewers

Why this role matters:

  • Your judgment directly shapes how the next generation of AI responds, reasons, and uses tools. This is high‑impact, flexible work for experts who want their depth to count, not faceless annotation work.Engagement model:Flexible hours, work on your own schedule
  • Long‑term engagements on frontier AI projects
  • Multiple delivery sprints across 3 to 6+ months
Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior AI Evaluation & RLHF Specialist
Senior AI Evaluation & RLHF Specialist

Innodata India • Dadri

In loco
INR 1.800.000 - 2.800.000
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Pune District

In loco
INR 1.858.273 - 2.106.043
Flexible remote work
Competitive pay up to $17/hour
Experience on advanced AI projects
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Ahmedabad District

In loco
INR 1.237.589 - 1.762.627
Flexible project schedule
Competitive hourly rates
Experience in advanced AI projects
AI Agent Evaluation Analyst (Freelance)
AI Agent Evaluation Analyst (Freelance)

Mindrift • Maharashtra

In loco
INR 1.242.200 - 1.490.640
Flexible schedule
Competitive pay up to $12/hour
Gain experience in advanced AI projects
AI Agent Evaluation Analyst
AI Agent Evaluation Analyst

Mindrift • New Delhi

In loco
INR 1.819.014
Flexible remote opportunity
Competitive hourly rates
Experience in advanced AI projects
Machine Learning Evaluation Specialist
Machine Learning Evaluation Specialist

Alignerr Corp. • Bengaluru

Remoto
INR 6.668.000 - 20.004.000
Remote work
Flexible schedule
Autonomy over schedule
+1
AI/ML Engineer (Contract)
AI/ML Engineer (Contract)

Deccan AI Experts • Hyderabad

In loco
INR 1.400.000 - 2.200.000
AI/ML Engineer (Contract)
AI/ML Engineer (Contract)

Deccan AI Experts • Bengaluru

In loco
INR 1.000.000 - 1.800.000
Staff AI Software Engineer
Staff AI Software Engineer

GE Vernova • Bengaluru

In loco
INR 3.000.000 - 5.400.000
Relocation assistance
Senior Research Scientist, Agent Evaluation
Senior Research Scientist, Agent Evaluation

Snow Planet • Hyderabad, Ahmedabad District

In loco
INR 2.400.000 - 4.200.000