RLHF Specialist

Odixcity Consulting

South Africa

Remote

ZAR 1,477,000 - 2,626,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Odixcity Consulting is seeking an RLHF Specialist to improve AI models through reinforcement learning from human feedback, focusing on data collection, evaluation, and alignment across a global team.

You will design prompts, create annotated datasets, and collaborate with ML engineers to enhance model safety, accuracy, and robust reasoning. Remote worldwide, full-time position with access to modern ML tooling.

Qualifications

  • Minimum of 2 years of experience in Data Annotation, Model Evaluation, Computational Linguistics, or Trust and Safety, specifically working with AI/ML training data.
  • Strong proficiency in Python and deep learning frameworks (PyTorch, JAX, or TensorFlow).
  • Deep understanding of Reinforcement Learning concepts (PPO, Trust Regions, Reward Hacking) and how they apply to language generation.
  • Hands-on experience fine-tuning open-source models (e.g., Llama 2/3, Mistral, gemma) using techniques like LoRA/QLoRA.
  • Experience working with annotation tools (LabelBox, Scale AI, Snorkel) and managing human-in-the-loop workflows.
  • Ability to diagnose why an RL policy collapsed and adjust hyperparameters or reward structure accordingly.
  • Experience with Constitutional AI or Self-Alignment techniques.
  • Contributions to open-source alignment libraries (TRL, Transformer Reinforcement Learning, Axolotl).
  • Experience with cloud Platforms (AWS SageMaker, GCP Vertex AI).

Responsibilities

  • Generate high-quality preference data by comparing multiple model responses and ranking them based on criteria such as helpfulness, honesty, and harmlessness (HHH).
  • Design complex, multi-turn prompts to stress-test model behavior and expose weaknesses in reasoning or safety.
  • Write detailed “chain-of-thought” explanations and rationales to train reward models on why specific responses are superior.
  • Collaborate with Machine Learning Engineers to analyze model failure modes and identify data gaps that, when filled, will improve reinforcement learning outcomes.
  • Develop and iterate on annotation strategies for preference scoring and reinforcement signals, ensuring consistency across a global team.
  • Proactively probe models to identify vulnerabilities, biases, or hallucination patterns, documenting findings for model optimization.
  • Analyze edge cases where the reward model behaves unexpectedly (e.g., over-indexing on verbosity or style over substance). Provide detailed feedback to ML engineers on reward model failure modes and suggest specific data interventions to correct model behavior.
  • Develop and document templated instruction sets for larger annotation teams. Translate complex reinforcement learning concepts into simple, repeatable tasks for junior reviewers, ensuring high-quality data collection at scale.
  • Monitor model performance over time by maintaining a personal test set of prompts. Regularly re-evaluate new model versions against historical benchmarks to track improvements or regressions in reasoning and alignment.

Skills

Python
Deep Learning
PyTorch
JAX
TensorFlow
Reinforcement Learning
PPO
LoRA/QLoRA
Open-source models
Annotation tools

Tools

LabelBox
Scale AI
Snorkel
AWS SageMaker
GCP Vertex AI

Job description

Job Title: RLHF Specialist

Location: Remote (Worldwide)

Job Summary: An RLHF Specialist is responsible for improving and aligning AI models using Reinforcement Learning from Human Feedback (RLHF) methodologies. This role focuses on designing, implementing, and optimizing feedback pipelines that enhance model performance, safety, factual accuracy, and alignment with human values.

Responsibilities
  • Generate high-quality preference data by comparing multiple model responses and ranking them based on criteria such as helpfulness, honesty, and harmlessness (HHH).
  • Design complex, multi-turn prompts to stress-test model behavior and expose weaknesses in reasoning or safety.
  • Write detailed “chain-of-thought” explanations and rationales to train reward models on why specific responses are superior.
  • Collaborate with Machine Learning Engineers to analyze model failure modes and identify data gaps that, when filled, will improve reinforcement learning outcomes.
  • Develop and iterate on annotation strategies for preference scoring and reinforcement signals, ensuring consistency across a global team.
  • Proactively probe models to identify vulnerabilities, biases, or hallucination patterns, documenting findings for model optimization.
  • Analyze edge cases where the reward model behaves unexpectedly (e.g., over-indexing on verbosity or style over substance). Provide detailed feedback to ML engineers on reward model failure modes and suggest specific data interventions to correct model behavior.
  • Develop and document templated instruction sets for larger annotation teams. Translate complex reinforcement learning concepts into simple, repeatable tasks for junior reviewers, ensuring high-quality data collection at scale.
  • Monitor model performance over time by maintaining a personal test set of prompts. Regularly re-evaluate new model versions against historical benchmarks to track improvements or regressions in reasoning and alignment.
Requirements
  • Minimum of 2 years of experience in Data Annotation, Model Evaluation, Computational Linguistics, or Trust and Safety, specifically working with AI/ML training data.
  • Strong proficiency in Python and deep learning frameworks (PyTorch, JAX, or TensorFlow).
  • Deep understanding of Reinforcement Learning concepts (PPO, Trust Regions, Reward Hacking) and how they apply to language generation.
  • Hands-on experience fine-tuning open-source models (e.g., Llama 2/3, Mistral, gemma) using techniques like LoRA/QLoRA.
  • Experience working with annotation tools (LabelBox, Scale AI, Snorkel) and managing human-in-the-loop workflows.
  • Ability to diagnose why an RL policy collapsed and adjust hyperparameters or reward structure accordingly.
  • Experience with Constitutional AI or Self-Alignment techniques.
  • Contributions to open-source alignment libraries (TRL, Transformer Reinforcement Learning, Axolotl).
  • Experience with cloud Platforms (AWS SageMaker, GCP Vertex AI).
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote RLHF Specialist: AI Alignment & Evaluation
Remote RLHF Specialist: AI Alignment & Evaluation

Odixcity Consulting • South Africa

Remote
ZAR 1,477,000 - 2,626,000
LLM Evaluator (Model Response Analyst)
LLM Evaluator (Model Response Analyst)

Odixcity Consulting • South Africa

Remote
ZAR 360,000 - 540,000
AI Engineering Lead
AI Engineering Lead

Network Finance • Randburg

On-site
ZAR 1,200,000 - 2,400,000
AI Engineer
AI Engineer

Vaimo • Pretoria

On-site
ZAR 600,000 - 800,000
Senior AI Engineer
Senior AI Engineer

PlaceTalent Pty • Gauteng

On-site
ZAR 900,000 - 1,200,000
Lead AI Engineer - NLP Specialist
Lead AI Engineer - NLP Specialist

Placements24 • Cape Town

Hybrid
ZAR 2,284,000 - 3,426,000
Fully remote work
Professional development budget
Access to cloud infrastructure and AIツ
+1
Lead AI Engineer - NLP Specialist
Lead AI Engineer - NLP Specialist

Placements24 • Stellenbosch

Hybrid
ZAR 2,332,000 - 3,498,000
Fully remote work
Flexible hours
Health, dental & vision coverage
+2
Lead AI Researcher - Machine Learning
Lead AI Researcher - Machine Learning

Placements24 • Sandton

Hybrid
ZAR 2,936,000 - 4,568,000
Competitive salary
Bonuses & equity
Remote-friendly benefits
+2
Mandarin Language Expert
Mandarin Language Expert

SME Careers • South Africa

Remote
ZAR 207,000 - 344,000
Ai Engineering Lead
Ai Engineering Lead

Network Recruitment • Johannesburg

On-site
ZAR 1,000,000 - 1,800,000