Machine Learning Researcher

Brahma Consulting Group

San Francisco (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Brahma Consulting Group is seeking a Research Scientist / Member of Technical Staff focused on Reinforcement Learning for large language models in Seattle. You will work hands-on at the frontier of RL for model alignment, balancing ideas, experiments, and real‑world deployment.

We value someone who can own ambiguous problems end to end, publish or demonstrate impact, and help the team advance RL methods such as post‑training, reward modeling, and related strategies for language models.

Qualifications

  • Strong foundations in RL and understanding why methods work.
  • Experience applying RL to language models and related post-training work.
  • A track record of research with measurable impact.

Responsibilities

  • Develop RL methods for large language models, including post-training and reward modeling.
  • Design and run scalable experiments and translate findings into reusable methods.
  • Tackle open problems and help decide what to pursue rather than follow fixed specs.

Skills

Reinforcement learning
Language models
Experiment design
RLHF/RLAIF

Job description

Research Scientist / Member of Technical Staff, Reinforcement Learning

We are running this search on behalf of a client building frontier models, and they need researchers who can push reinforcement learning forward at the level of language models, not just apply it. This is hands‑on research with a tight loop between ideas, experiments, and what ships into real systems.

What you will work on
  • Reinforcement learning methods for large language models: post‑training, reward modeling, and the algorithms that make models more capable and more aligned with what we actually want from them.
  • Designing and running experiments at scale, then turning what you learn into methods the rest of the team can build on.
  • Open problems rather than closed tickets. You will help decide what is worth working on, not just execute someone else's spec.
Who fits
  • Strong foundations in reinforcement learning, including the core ideas that predate the current wave. You understand why methods work, not only how to call them.
  • Direct experience applying RL to language models. This could be RLHF, RLAIF, preference optimization, reward modeling, or related post‑training work.
  • A track record of research you can point to: published papers, strong open‑source contributions, or shipped methods with measurable impact.
  • Comfort owning ambiguous problems end to end, from framing the question to landing the result.
Nice to have
  • Experience training or post‑training models at scale.
  • A history of moving quickly from idea to working experiment.
  • Strong written communication. We share results across the team often.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Research Scientist — Large Language Models
RL Research Scientist — Large Language Models

Brahma Consulting Group • Seattle (WA)

On-site
USD 150,000 - 230,000
Member of Technical Staff, Reinforcement Learning
Member of Technical Staff, Reinforcement Learning

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Machine Learning Researcher
Machine Learning Researcher

Thurn Partners • Miami (FL)

On-site
USD 180,000 - 240,000
Research Engineer
Research Engineer

Nace.AI • Palo Alto (CA)

On-site
USD 120,000 - 160,000
RL Researcher, Product Systems
RL Researcher, Product Systems

Autohand AI Ltd. • San Francisco (CA)

On-site
USD 150,000 - 210,000
Publications
RL Research Engineer - Scalable, Safe AI Systems
RL Research Engineer - Scalable, Safe AI Systems

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2
Research Scientist - Post-training / RL
Research Scientist - Post-training / RL

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Scientist - Reinforcement Learning
Research Scientist - Reinforcement Learning

Optimized, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 185,000 - 255,000