As an AI/ML Research Engineer – LLM Post-Training & Evaluation, you will design and build the pipelines and tooling that connect data, evaluation, and model post-training. You’ll work on areas such as SFT, DPO/preference optimization, RLHF/RLAIF workflows, automated evaluation, experiment tracking, and multimodal model assessment.
This is a hands-on engineering role suited for someone who can bridge research and production-quality ML engineering.
Key Responsibilities
- Lead or co-lead technically complex ML engineering projects from discussion through delivery
- Build and optimize LLM training and post-training pipelines
- Develop data ingestion, preprocessing, fine-tuning, evaluation, and experiment-tracking workflows
- Build automated evaluation pipelines, benchmarks, and task-specific test harnesses
- Integrate human and AI-augmented evaluation signals into model development
- Improve reproducibility, metrics logging, regression monitoring, and experiment reliability
- Diagnose model behavior, training issues, data problems, and evaluation inconsistencies
- Work with Language Data Scientists and Applied Research Scientists to implement evaluation frameworks
- Collaborate directly with customer technical stakeholders
- Contribute to internal R&D, benchmark frameworks, evaluation tooling, and reusable ML infrastructure
- Mentor junior engineers and contribute to technical design reviews and engineering standards
Must-Have Qualifications
- 2–3+ years of relevant ML engineering, applied ML systems, or research engineering experience
- BS/MS/PhD in Computer Science, Machine Learning, AI, Applied Mathematics, or a related quantitative technical field
- Hands-on experience with LLM training, fine-tuning, or post-training
- Experience with at least one of:
- Supervised Fine-Tuning (SFT)
- Preference Optimization such as DPO
- Foundation model/domain adaptation
- Strong Python programming and software engineering fundamentals
- Experience with modern ML frameworks such as PyTorch, JAX, or TensorFlow
- Experience with model libraries/tooling such as the Hugging Face ecosystem, vLLM, or distributed training stacks
- Experience designing LLM/ML evaluation pipelines, including metrics, datasets, and experiment comparisons
- Understanding of ML systems engineering, reproducibility, observability, and debugging
- Experience with distributed ML systems and performance optimization, preferably in GPU/accelerator environments
- Experience with large-scale data processing and workflow orchestration supporting ML workloads
- Ability to collaborate with research scientists, ML engineers, data engineers, and technical stakeholders
Good-to-Have
- Multimodal model training/evaluation involving text, image, audio, or video
- Long-context evaluation or model adaptation
- Agentic or multi-turn evaluation, tool-use simulation, or interactive environment testing
- Customer-facing technical consulting, solutions engineering, or applied research delivery
- LLM safety, alignment, robustness, or red-teaming evaluation experience
- Open-source ML/LLM contributions or relevant technical publications
Why Join Us
- Work on advanced GenAI systems involving LLM training, post-training, and evaluation
- Bridge research and engineering by turning evaluation findings into measurable model improvements
- Collaborate with AI experts across research, data science, engineering, and technical delivery
- Contribute to R&D including benchmarks, evaluation frameworks, and reusable ML infrastructure