Machine Learning Engineer — Self-Improving RL Pipelines

Hippocratic AI Inc.

Menlo Park (CA)

On-site

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Hippocratic AI in Menlo Park, CA is seeking an engineering-focused ML role to build and maintain end-to-end self-improvement loops with strong emphasis on reproducibility and safety. You will design reward signals, evaluation harnesses, and data pipelines for robust, production-ready systems.

Ideal candidates have hands-on RLHF/RLAIF experience, expertise in large-scale training, and a track record of shipping ML systems that stay healthy over time.

Qualifications

  • Excellent Python and clean, well-tested ML training code.
  • Solid grasp of data pipelines, distributed / large-scale training, and experiment tracking.
  • The instinct and skill to debug why a model silently got worse — not just why it crashed.
  • Hands-on experience with a feedback or learning loop (RLHF/RLAIF, active‑learning, or data flywheels).
  • Experience retraining or continual-learning pipelines using production data or model outputs.

Responsibilities

  • Build and maintain the training, evaluation, and deployment loops at the core of the self‑improvement system, with emphasis on reproducibility and reliability.
  • Design and implement reward and feedback signals; mitigate reward hacking, specification gaming, and distribution drift.
  • Build evaluation harnesses and metrics before models to measure improvement.
  • Own data pipelines and automated data flywheels feeding the learning loop.
  • Debug subtle model-quality regressions and stabilize non‑stationary training and feedback loops.
  • Collaborate with research and product to turn methods into robust, shippable systems.

Skills

Python
ML training code
Data pipelines
Distributed training
Experiment tracking
Debugging models
RLHF / RLAIF

Education

PhD in RL/ML
MS in RL/ML

Tools

LLM fine-tuning tooling
Agent orchestration tools

Job description

Hippocratic AI in Menlo Park, CA is seeking an engineering-focused ML role to build and maintain end-to-end self-improvement loops with strong emphasis on reproducibility and safety. You will design reward signals, evaluation harnesses, and data pipelines for robust, production-ready systems.

Ideal candidates have hands-on RLHF/RLAIF experience, expertise in large-scale training, and a track record of shipping ML systems that stay healthy over time.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Applied Scientist — Healthcare RL & Safe AI
Staff Applied Scientist — Healthcare RL & Safe AI

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 230,000 - 290,000
Machine Learning Engineer
Machine Learning Engineer

Hippocratic AI Inc. • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Staff RL Scientist - Healthcare AI
Staff RL Scientist - Healthcare AI

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)
Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 240,000
ML Systems Engineer — RL Training & Finetuning
ML Systems Engineer — RL Training & Finetuning

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+1
Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)
Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 230,000 - 290,000
Staff ML Engineer — RL & Production Systems
Staff ML Engineer — RL & Production Systems

People In AI • San Francisco (CA)

Hybrid
USD 270,000 - 280,000
Staff RL Data Platform Engineer — Full-Stack & Data Pipelines
Staff RL Data Platform Engineer — Full-Stack & Data Pipelines

Anthropic Limited • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 405,000
Equity donation matching
Generous vacation & parental leave
Flexible working hours
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior ML Engineer: Learned Driving & RL Pipelines
Senior ML Engineer: Learned Driving & RL Pipelines

PlusAI • Santa Clara (CA)

On-site
USD 130,000 - 220,000
Catered lunch
Unlimited snacks
401(k) plan