Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Hippocratic AI is seeking experts to own the Reinforcement Learning (RL) and On-Policy Distillation (OPD) post-training pipeline for healthcare LLMs. You will improve clinical reasoning, safety, and alignment, deploying models to interact with millions of patients across diverse clinical use cases.
The role requires 5+ years in NLP, LLM training, or RL, with 2+ years in RL for LLM post-training, and hands-on experience with large-scale multi-node LLMs in a Palo Alto setting.
Hippocratic AI is seeking experts to own the Reinforcement Learning (RL) and On-Policy Distillation (OPD) post-training pipeline for healthcare LLMs. You will improve clinical reasoning, safety, and alignment, deploying models to interact with millions of patients across diverse clinical use cases.
The role requires 5+ years in NLP, LLM training, or RL, with 2+ years in RL for LLM post-training, and hands-on experience with large-scale multi-node LLMs in a Palo Alto setting.