The Principal Machine Learning Engineer in Applied AI focuses on turning advanced AI models into dependable, customer-ready workflows. The work emphasizes model post-training, evaluation, and production-oriented ML systems, with an end goal of reliable use in customer scientific contexts.
Responsibilities
- Use deep, hands-on expertise to address the last-mile gap between model capability and customer-specific scientific workflows.
- Lead post-training efforts using methods such as SFT and RL (including DPO and PPO/GRPO) to align model behavior with customer requirements and feedback.
- Design and run evaluation loops to measure model quality, reliability, and fit for customer use cases, applying patterns refined across many prior projects.
- Translate customer learnings, data signals, and evaluation outcomes into concrete model improvement cycles.
- Collaborate with AI researchers to convert model advances into reliable, usable capabilities.
- Work with Software teams to integrate model behavior into end-to-end product workflows.
- Debug complex model failures using traces, evaluations, customer context, and scientific feedback.
- Mentor engineers and share reusable patterns for model adaptation, evaluation, and deployment, based on experience shipping ML systems.
Requirements
- Minimum 2-3 years of hands-on post-training experience, including SFT and RL methods such as DPO/PPO/GRPO, along with evaluation system design developed through repeated problem-solving.
- Strong software engineering skills in Python and modern ML frameworks such as PyTorch.
- Proven ability to debug ambiguous, high-stakes model behavior quickly by using data, traces, logs, and qualitative feedback, informed by many prior failure modes.
- Experience leading technical work across research and engineering teams.
- Deep familiarity with large language models, multi-modal models, or agentic AI systems.
- Clear communication skills to translate customer needs into technical approaches and explain complex model behavior to technical and non-technical audiences.
Technologies
- Python, PyTorch
- SFT
- RL, DPO, PPO, GRPO
- RLHF
- MoE
Bonus Points
- Experience adapting models for customer-facing or production workflows, preferably in scientific, technical, or data-intensive domains.
- Experience with RL post-training such as RLHF, GRPO, or tool-augmented RL.
- Experience building evaluation harnesses, model monitoring, or quality dashboards.
- Experience training MoE architectures.
- A history of mentoring engineers and serving as a go-to technical expert, rather than a people manager or strategy owner.
Location and Compensation
Location: Cambridge, MA (onsite).
Salary: USD 252,000 - 336,000 per year.