Domain Expert - Mathematics

Innodata Inc.

India

On-site

INR 900,000 - 1,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Innodata Inc. is seeking Reward Validation Specialists to contribute to advanced AI training and evaluation projects. PhD-qualified researchers with RL, optimization, and Python skills will help improve the reliability and accuracy of AI evaluation systems.

You will build and validate automated grading systems for AI model evaluation pipelines, focusing on step-by-step reward logic and scalable Python tools to ensure consistent, objective assessments across large workloads.

Qualifications

  • PhD or PhD candidate nearing completion in Computer Science, ML, AI, Electrical Engineering, Applied Mathematics, Robotics or a similar quantitative field.
  • Strong understanding of reinforcement learning concepts, including MDPs, value functions, credit assignment, and reward shaping.
  • Solid foundation in optimization techniques, including gradient-based and convex/non-convex optimization.
  • Strong Python programming skills with experience writing, testing, and debugging code.
  • Ability to quickly learn and independently apply new evaluation methodologies.

Responsibilities

  • Validate step-level reward logic to ensure it accurately measures the intended model behavior.
  • Verify that grading criteria align with task instructions and evaluation objectives.
  • Ensure models have the required context, files, and information needed to complete each evaluated step.
  • Develop, implement, and test automated graders using Python.
  • Analyze and improve reward functions to support scalable and reliable AI evaluation pipelines.
  • Collaborate with internal teams to refine evaluation methodologies and maintain grading consistency.

Skills

Reinforcement Learning
Python programming
Optimization
RL concepts (MDP, value functions)

Education

PhD in Computer Science / ML / AI / EE / Applied Math

Tools

NumPy
PyTorch
Python

Job description

We are hiring Reward Validation Specialists to contribute to advanced AI training and evaluation projects. This role is ideal for PhD-qualified researchers with expertise in reinforcement learning, optimization, and Python programming who are passionate about improving the reliability and accuracy of AI evaluation systems.

Role Overview

In this role, you will help build and validate automated grading systems for AI model evaluation pipelines. Rather than assessing only the final output, you will evaluate and verify step-by-step reward logic to ensure each grading component accurately measures the intended behavior. You will also develop and implement automated graders in Python that can scale across large evaluation workflows while maintaining consistency, accuracy, and alignment with task requirements.

Key Responsibilities
  • Validate step-level reward logic to ensure it accurately measures the intended model behavior.
  • Verify that grading criteria align with task instructions and evaluation objectives.
  • Ensure models have the required context, files, and information needed to complete each evaluated step.
  • Develop, implement, and test automated graders using Python.
  • Analyze and improve reward functions to support scalable and reliable AI evaluation pipelines.
  • Collaborate with internal teams to refine evaluation methodologies and maintain grading consistency
Mandatory Requirements
  • PhD (or PhD candidate nearing completion) in Computer Science, Machine Learning, Artificial Intelligence, Electrical Engineering, Applied Mathematics, Robotics, or another quantitative discipline.
  • Strong understanding of reinforcement learning concepts, including Markov Decision Processes (MDPs), value functions, credit assignment, and reward shaping.
  • Solid foundation in optimization techniques, including gradient-based methods and convex/non-convex optimization.
  • Strong Python programming skills with experience writing, testing, and debugging code.
  • Ability to quickly learn and independently apply new evaluation methodologies.
Preferred Qualifications
  • Experience with control theory, dynamic programming, optimal control, or stability analysis.
  • Familiarity with RLHF (Reinforcement Learning from Human Feedback), process reward models, or AI evaluation frameworks.
  • Hands-on experience with NumPy, PyTorch, or other machine learning libraries.
  • Background in machine learning, reinforcement learning, robotics, control systems, or applied mathematics.
  • Experience developing evaluation pipelines, automated testing frameworks, or model validation tools.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Science Expert - Evaluation Specialist
Data Science Expert - Evaluation Specialist

Mercor • Mumbai

On-site
INR 2,500,000 - 4,000,000
Data Science Expert - Evaluation Specialist
Data Science Expert - Evaluation Specialist

Obsidian • Mumbai

On-site
INR 1,800,000 - 3,000,000
Mathematics Expert (PhD) - 34430
Mathematics Expert (PhD) - 34430

Turing • Hyderabad

Remote
Data Science Specialist - Fully Remote
Data Science Specialist - Fully Remote

Mercor • Mumbai

Remote
INR 2,400,000 - 3,600,000
Remote Maths AI Trainer (PhD) - 34430
Remote Maths AI Trainer (PhD) - 34430

Turing • Mumbai Suburban

Remote
INR 5,273,000 - 11,864,000
Work on cutting-edge AI research
Collaborate with global experts
Flexible hours
AI Researcher
AI Researcher

Provue • Mumbai

On-site
INR 1,200,000 - 1,800,000
AI Assessment Specialist - PhD
AI Assessment Specialist - PhD

Mercor • Mumbai

On-site
INR 6,686,000 - 10,506,000
Sr. Responsible AI Analyst
Sr. Responsible AI Analyst

Infosys • Bengaluru

On-site
INR 2,500,000 - 3,500,000
Senior AI / ML engineer
Senior AI / ML engineer

Srs Business Solutions India • Hyderabad

On-site
INR 2,500,000 - 3,500,000
Computational Statistics Expert - PhD
Computational Statistics Expert - PhD

Mercor • Mumbai

On-site
INR 400,000 - 700,000