Senior AI/LLM Engineer

Innodata Inc.

Philippines

On-site

PHP 2,400,000 - 4,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Innodata Inc. is seeking a Senior AI/LLM Engineer to lead training, alignment, and optimization of large language models. You will own the full post-training pipeline from supervised fine-tuning through reward modeling and RL optimization, ensuring efficient production deployment.

You will design SFT pipelines for open-weight models, build reward models, and implement policy-gradient techniques. Strong code skills in Python, C/C++, and Rust are required for performance-critical paths.

Qualifications

  • Bachelor’s degree in IT, CS, Engineering or related field.
  • 6+ years of professional software/ML engineering experience.
  • 3+ years hands-on with production ML/DL systems and 2+ years building with LLMs; production shipping experience.

Responsibilities

  • Own and drive full RLHF pipeline: data collection, reward model training, RL fine-tuning.
  • Design and run SFT pipelines on open-weight models as foundation for RLHF.
  • Build and train reward models that capture human preferences.
  • Implement Connsitutional AI and RLAIF techniques to reduce annotation needs.
  • Write and tune low-level C/C++ and Rust code for inference performance.

Skills

Python
C/C++
Rust
RLHF
LLaMA/Mistral/Qwen
Distributed training
Transformer architectures

Education

Bachelor’s degree in IT / CS / Engineering

Tools

llama.cpp
vLLM
TensorRT
LoRA/QLoRA
FAISS/Milvus

Job description

As a Senior AI/LLM Engineer, you will lead our efforts to train, align, and optimize large language models. You will own the full post-training pipeline from supervised fine-tuning through reward modeling and RL optimization, while also ensuring models run efficiently in production. This is a role that bridges alignment research and systems engineering.

What You’ll Own
  • Own and drive the full RLHF pipeline: data collection, reward model training, and RL fine-tuning using PPO, DPO, GRPO, and RLAIF
  • Design and run Supervised Fine-Tuning (SFT) pipelines on open-weight models (LLaMA, Mistral, Qwen) as the foundation for RLHF
  • Build and train reward models that accurately capture human preferences from annotation data
  • Design human feedback collection pipelines: labeling rubrics, annotator calibration, and preference dataset curation
  • Implement Constitutional AI and RLAIF techniques to reduce reliance on costly human annotation
  • Red team models post-training — probing for jailbreaks, regressions, unsafe outputs, and alignment failures
  • Design and maintain evaluation benchmarks to measure alignment, safety, and capability before and after RL training
  • Optimize inference pipelines and runtimes (llama.cpp, vLLM, TensorRT) to serve aligned models efficiently at scale
  • Implement quantization strategies (INT4/INT8/FP8, LoRA, QLoRA) to deploy fine-tuned models on target hardware
  • Write and tune low-level C/C++ and Rust code for inference performance where Python cannot reach
  • Diagnose and resolve training instabilities, reward hacking, and production inference bugs under pressure
  • Stay at the frontier — read alignment and RL papers weekly and translate findings into working experiments
What You’ll Bring
  • Bachelor’s degree in IT, Computer Science, Engineering or a related field
  • 6+ years of professional software/ML engineering experience
  • 3+ years hands-on with production ML/DL systems and 2+ years building with LLMs; a proven track record of shipping AI/ML systems to production at scale.
  • 2+ years in a tech lead (or equivalent) role on AI/ML projects — leading the technical delivery of a team or workstream, owning architecture decisions, running design reviews, and mentoring engineers (with or without formal people-management responsibility).
What You’ll Bring
  • Deep familiarity with policy gradient methods: PPO stability, KL divergence constraints, reward shaping
  • Experience with Direct Preference Optimization (DPO) and its variants as an RLHF alternative
  • Understanding of reward hacking, Goodhart’s Law, and mitigation strategies in RL training
  • Familiarity with RLAIF (RL from AI Feedback) and Constitutional AI approaches
  • Ability to design preference datasets and annotation rubrics that produce high-quality reward signal
  • Experience diagnosing training instabilities: reward collapse, mode collapse, KL divergence blowup
  • Python as the primary language for all training, fine-tuning, and evaluation pipelines
  • Strong mathematical foundation: RL theory, probability, linear algebra, optimization — deep enough to derive loss functions and debug training dynamics
  • C and C++ for systems-level inference work, runtime contributions, and performance-critical paths
  • Rust experience with ML tooling.
  • Familiarity with transformer architecture, attention, tokenization, and how post-training interacts with pretraining
  • Experience with distributed training frameworks for large-scale fine-tuning
  • Experience with vector databases such as FAISS or Milvus
  • Familiarity with retrieval-augmented generation (RAG) pipelines
  • Experience integrating LLMs with external tools, APIs, and agent-based systems
  • Exposure to Rapid Application Development (RAD) approaches for building and iterating AI solutions efficiently
Our Commitment to Diversity, Equity, Inclusion & Belonging

At Innodata, we believe the best teams are built from a wide range of backgrounds, perspectives, and lived experiences — and that diverse teams build better data, better products, and a better workplace. We are committed to fostering an environment where everyone feels welcomed, respected, valued, and empowered to do their best work.

We are an equal opportunity employer. We make all employment decisions on the basis of merit, qualifications, and business need, and we do not discriminate on the basis of race, ethnicity, color, religion, sex, gender identity or expression, sexual orientation, age, disability, marital or family status, national origin, or any other characteristic protected by applicable law.

We are committed to building an inclusive and accessible hiring process.

If you require any accommodation to participate fully in our recruitment process, please let us know; we're glad to support you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/LLM Team Manager – Expression of Interest
AI/LLM Team Manager – Expression of Interest

Innodata Inc. • Philippines

On-site
PHP 900,000 - 1,500,000
Senior AI/LLM Engineer: RLHF & Production Lead
Senior AI/LLM Engineer: RLHF & Production Lead

Innodata Inc. • Philippines

On-site
PHP 2,400,000 - 4,800,000
Head of Engineering, AI/LLM Practice (Associate VP, Delivery Engineering)
Head of Engineering, AI/LLM Practice (Associate VP, Delivery Engineering)

Innodata Knowledge Services, Inc. • Philippines

On-site
PHP 900,000 - 1,500,000
Flexible work arrangements
DEIB commitment and programs
Equal opportunity employer
AI/LLM Software Engineer - Expression of Interest
AI/LLM Software Engineer - Expression of Interest

Innodata Inc. • Philippines

Hybrid
PHP 700,000 - 1,200,000
Machine Learning Engineer
Machine Learning Engineer

Nezda Global • Philippines

On-site
PHP 600,000 - 1,200,000
Senior Software Engineering Lead
Senior Software Engineering Lead

Vanigent Biopharm • Metro Manila

On-site
PHP 3,000,000 - 5,000,000
AI Developer – Backend & LLM Systems
AI Developer – Backend & LLM Systems

Salvo Software LLC • Mexico

On-site
PHP 1,228,000 - 2,150,000
Senior Python AI Engineer
Senior Python AI Engineer

Proxify • Mexico

On-site
MXN 1,380,000 - 1,899,000
Guaranteed on‑time monthly payments
Up to 24 flex days off per year
Career-accelerating opportunities
+1
Senior Training and Quality Manager
Senior Training and Quality Manager

Innodata Inc. • Philippines

On-site
PHP 900,000 - 1,500,000
Diversity, Equity, Inclusion & Belongi
Flexible work arrangements
Large Language Model Algorithm Engineer
Large Language Model Algorithm Engineer

Binance • Hinoba-an

On-site
PHP 1,200,000 - 2,400,000