Staff MLE, Reinforcement Learning

People In AI

San Francisco (CA)

Hybrid

USD 270,000 - 280,000

Full time

47 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

People In AI in San Francisco is seeking a Staff-level Machine Learning Engineer who blends deep reinforcement learning with production software engineering to shape experiments, model improvements, and ML infrastructure.

You will work between research and production, set technical direction, lead other engineers, and design systems for large-scale training and evaluation while remaining hands-on.

Qualifications

  • Strong hands-on ML engineering experience and production mindset.
  • Experience with model post-training or fine-tuning.
  • Experience with RL or preference-optimization approaches such as GRPO, PPO, DPO, or similar.

Responsibilities

  • Design and build reinforcement learning environments for agentic tasks.
  • Develop task definitions, tool interfaces, reward structures, state management, and evaluation logic.
  • Build post-training and fine-tuning pipelines across supervised fine-tuning and reinforcement learning.
  • Develop verifiers, graders, rubrics, and evaluation systems for complex model behavior.
  • Run and diagnose model-training experiments and ensure data and reward quality.

Skills

Hands-on ML engineering
Python
ML infra / distributed systems
System design
Technical leadership
Staff-level influence
Technical communication

Tools

Reinforcement Learning frameworks
Model evaluation tools

Job description

Compensation: $270,000 - $280,000 base + equity

Location: San Francisco, hybrid 3 days per week

Join a fast-growing AI technology company building the infrastructure, training environments, and evaluation systems used to improve advanced AI models.

This is a Staff-level role for a Machine Learning Engineer who combines real depth in reinforcement learning and post-training with strong production software engineering. The company is looking for someone who can operate across experimentation, model improvement, infrastructure, and technical leadership while remaining deeply hands-on.

The Mission

The company is building systems that help advanced AI models learn, improve, and perform reliably on increasingly complex tasks.

That means creating reinforcement learning environments, generating high-quality training signal, evaluating model behavior, building reliable graders and verifiers, and developing the infrastructure required to run large-scale training and evaluation workflows.

The work sits much closer to the underlying models and training lifecycle than traditional AI application development.

The Role

You will sit between research engineering and production Machine Learning Engineering, combining hands-on experimentation with Staff-level technical ownership.

You will work on problems across reinforcement learning, post-training, agent training, evaluation, and ML infrastructure. Projects move quickly, and you may move between running experiments, designing systems, writing production code, setting technical direction, and leading other engineers through ambiguous technical problems.

The existing team has strong implementers. This hire is intended to bring another level of technical judgment, helping determine what should be built, how it should be designed, and how the team should execute.

What You'll Do
  • Design and build reinforcement learning environments for agentic tasks.
  • Develop task definitions, tool interfaces, reward structures, state management, and evaluation logic.
  • Build post-training and fine-tuning pipelines across supervised fine-tuning and reinforcement learning.
  • Develop verifiers, graders, rubrics, and evaluation systems for complex model behavior.
  • Run and diagnose model-training experiments, including issues around reward quality, data quality, and training signal.
  • Build infrastructure capable of running large numbers of model and agent trajectories.
  • Develop production-grade ML systems across orchestration, reliability, fault tolerance, and experiment management.
  • Translate ambiguous technical problems into clear architectures and execution plans.
  • Set technical direction and influence other engineers while remaining deeply hands-on.
  • Use modern AI development tools while maintaining strong engineering judgment around the resulting systems.
What You'll Bring
  • Strong hands-on Machine Learning Engineering experience.
  • Practical experience with model post-training or fine-tuning.
  • Experience with SFT and at least one RL or preference-optimization approach such as GRPO, PPO, DPO, or similar.
  • Experience with agent environments, model evaluation, reward design, verifiers, graders, or adjacent areas.
  • Strong Python skills and production software engineering fundamentals.
  • Experience with ML infrastructure, distributed systems, platform engineering, or data systems at scale.
  • Strong system design and architecture judgment.
  • Ability to diagnose why a model or training run is or is not improving.
  • Evidence of Staff-level technical leadership and influence across other engineers.
  • High agency and a track record of independently identifying important technical problems and driving them through to completion.
  • Comfort working in a fast-moving, ambiguous engineering environment.
  • Strong technical communication skills.
Why Join?
  • Work directly on reinforcement learning, post-training, agent evaluation, and advanced ML infrastructure.
  • Operate closer to the underlying model-development lifecycle than traditional AI application engineering.
  • Combine research-oriented ML problems with real production engineering responsibility.
  • Stay deeply hands-on while having meaningful Staff-level influence over architecture and technical direction.
  • Work across a broad range of rapidly evolving AI problems rather than being siloed into one narrow technical area.
  • Join an engineering culture that values technical judgment, ownership, speed, and individual impact.
  • Build systems focused on measurable model improvement rather than isolated demos or API integrations.
About People In AI

We partner with AI-first startups, scale-ups, and enterprise organizations to connect exceptional engineers with opportunities to build production AI systems, intelligent platforms, and the next generation of AI infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Applied AI
Research Engineer, Applied AI

HeyMilo AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Staff Software Engineer, AI
Staff Software Engineer, AI

MEA Energy • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 300,000
Staff Software Engineer, AI
Staff Software Engineer, AI

International Association of Plumbing and Mechanical Officials (IAPMO) • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 300,000
Staff ML Engineer — RL & Production Systems
Staff ML Engineer — RL & Production Systems

People In AI • San Francisco (CA)

Hybrid
USD 270,000 - 280,000
Machine Learning Engineer
Machine Learning Engineer

Harnham • San Francisco (CA)

Hybrid
USD 250,000 - 350,000
Equity
Full benefits
Member of Technical Staff - Machine Learning Capabilities, New Graduates
Member of Technical Staff - Machine Learning Capabilities, New Graduates

Preference Model • Seattle (WA)

On-site
USD 165,000 - 200,000
Competitive cash and equity compensation
Health, vision, dental benefits
401K match
+3
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Recruiting From Scratch • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Competitive equity
Cutting-edge diffusion models
Direct collaboration with researchers
Senior AI/ML Engineer
Senior AI/ML Engineer

CB Smart Recruit • Los Angeles (CA)

On-site
USD 180,000 - 350,000
Competitive sign-on bonus
Comprehensive benefits package
Member of Technical Staff - Machine Learning Capabilities
Member of Technical Staff - Machine Learning Capabilities

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Health, vision, dental
401K match
Lunch onsite
+3
Machine Learning Engineer
Machine Learning Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 150,000 - 190,000