Senior AI/ML Developer

Innodata India Private Limited

Dadri

On-site

INR 1,800,000 - 3,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Innodata India Private Limited is seeking an experienced AI/ML engineer to spearhead RLHF pipeline development, reward model design, and fine-tuning of open-weight models. You will build SFT pipelines, design annotation rubrics, and advance RLAIF/Constitutional AI approaches while optimizing inference performance at scale.

The role requires deep knowledge of RLHF mechanics, Python for ML, and systems-level coding in C/C++ and Rust, with hands-on experience in distributed training and vector

Qualifications

  • Hands-on RLHF end-to-end experience.

Responsibilities

  • Own and drive the full RLHF pipeline: data collection, reward model training, and RL fine-tuning.

Skills

RLHF end-to-end
Policy gradient methods
PPO stability
DPO variants
Reward shaping
RLAIF
Constitutional AI
Annotation rubric design
Python for ML
C/C++ systems programming
Rust ML tooling
Transformer architectures
Distributed training
Vector databases
RAG pipelines
API/integration

Tools

llama.cpp
vLLM
TensorRT
LoRA
QLoRA
FAISS
Milvus

Job description

Responsibilities
  • Own and drive the full RLHF pipeline: data collection, reward model training, and RL fine‑tuning using PPO, DPO, GRPO, and RLAIF
  • Design and run Supervised Fine‑Tuning (SFT) pipelines on open-weight models (LLaMA, Mistral, Qwen) as the foundation for RLHF
  • Build and train reward models that accurately capture human preferences from annotation data
  • Design human feedback collection pipelines: labeling rubrics, annotator calibration, and preference dataset curation
  • Implement Constitutional AI and RLAIF techniques to reduce reliance on costly human annotation
  • Red team models post‑training — probing for jailbreaks, regressions, unsafe outputs, and alignment failures
  • Design and maintain evaluation benchmarks to measure alignment, safety, and capability before and after RL training
  • Optimize inference pipelines and runtimes (llama.cpp, vLLM, TensorRT) to serve aligned models efficiently at scale
  • Implement quantization strategies (INT4/INT8/FP8, LoRA, QLoRA) to deploy fine‑tuned models on target hardware
  • Write and tune low‑level C/C++ and Rust code for inference performance where Python cannot reach
  • Diagnose and resolve training instabilities, reward hacking, and production inference bugs under pressure
  • Stay at the frontier — read alignment and RL papers weekly and translate findings into working experiments
Core Requirements And Technical Skills
  • Hands‑on experience implementing RLHF end‑to‑end — not just using libraries, but understanding the mechanics
  • Deep familiarity with policy gradient methods: PPO stability, KL divergence constraints, reward shaping
  • Experience with Direct Preference Optimization (DPO) and its variants as an RLHF alternative
  • Understanding of reward hacking, Goodhart’s Law, and mitigation strategies in RL training
  • Familiarity with RLAIF (RL from AI Feedback) and Constitutional AI approaches
  • Ability to design preference datasets and annotation rubrics that produce high‑quality reward signal
  • Experience diagnosing training instabilities: reward collapse, mode collapse, KL divergence blowup
  • Python as the primary language for all training, fine‑tuning, and evaluation pipelines
  • Strong mathematical foundation: RL theory, probability, linear algebra, optimization — deep enough to derive loss functions and debug training dynamics
  • C and C++ for systems‑level inference work, runtime contributions, and performance‑critical paths
  • Rust experience with ML tooling.
  • Familiarity with transformer architecture, attention, tokenization, and how post‑training interacts with pretraining
  • Experience with distributed training frameworks for large‑scale fine‑tuning
  • Experience with vector databases such as FAISS or Milvus
  • Familiarity with retrieval‑augmented generation (RAG) pipelines
  • Experience integrating LLMs with external tools, APIs, and agent‑based systems
  • Exposure to Rapid Application Development (RAD) approaches for building and iterating AI solutions efficiently
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Evaluation & RLHF Specialist
Senior AI Evaluation & RLHF Specialist

Innodata India • Dadri

On-site
INR 1,800,000 - 2,800,000
Research Engineer, Multi-Domain Alignment (SLM)
Research Engineer, Multi-Domain Alignment (SLM)

Indian AI Research Organisation (IAIRO) • India

On-site
INR 1,800,000 - 3,200,000
AI Developer
AI Developer

Salvo Software LLC • India

On-site
INR 900,000 - 1,500,000
AI Engineer
AI Engineer

Aziro • Chennai District, Bengaluru, Pune District

Hybrid
INR 1,200,000 - 1,800,000
Founding AI Research Lead
Founding AI Research Lead

LH2 AI Labs • Bengaluru

On-site
INR 400,000 - 700,000
Principal AI/ML Engineer
Principal AI/ML Engineer

Optum India • Bengaluru

On-site
INR 4,000,000 - 8,000,000
AI/ML Engineer
AI/ML Engineer

Creuto Cloud Private Limited • India

On-site
INR 800,000 - 1,600,000
RL Environment Researcher
RL Environment Researcher

Provue • Mumbai

On-site
INR 1,500,000 - 2,800,000
Prismforce Pvt Ltd - AI Engineer - LLM/RAG
Prismforce Pvt Ltd - AI Engineer - LLM/RAG

Prismforce • Maharashtra

On-site
INR 1,800,000 - 3,200,000
Backend AI Engineer
Backend AI Engineer

CirrusLabs • Bengaluru

On-site
INR 4,000,000 - 7,000,000