RL Research Scientist - Post-Training on LLMs & Code Models

AMD

Santa Clara (CA)

On-site

USD 150,000 - 230,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Benefits at a glance

Job summary

AMD is hiring an AI Research Scientist specializing in reinforcement learning for post-training and interactive learning of large generative models applied to engineering and hardware tasks. You will design RL algorithms, run studies, and collaborate with infra and product teams to land methods that improve task success and stability.

The role requires a PhD (preferred) and a strong publication record, with hands-on experience training RL or preference-based models at scale, especially with LLMs

Qualifications

  • PhD in Computer Science, Machine Learning, or related field strongly preferred.
  • Strong publication record in reinforcement learning.
  • Hands‑on experience training RL or preference‑based models at non‑trivial scale (GPUs, distributed jobs).
  • Experience with LLM post‑training, RLHF/RLAIF, or policy optimization for language or code agents.

Responsibilities

  • Research and develop RL methods for post‑training LLMs and code models on structured engineering tasks with verifiable or preference‑based feedback
  • Design reward models, curricula, and off‑policy or on‑policy training recipes suited to sparse, noisy, or expensive labels from experts and simulators
  • Characterize failure modes (reward hacking, degenerate policies, instability) and propose mitigations grounded in experiments
  • Collaborate with RL infra engineers to scale training; define interfaces for rollout generation, logging, and reproducibility
  • Publish at top venues (e.g. NeurIPS, ICML, ICLR) and contribute internal technical leadership on the RL roadmap

Skills

RL theory
RL research
RLHF/RLAIF

Education

PhD in CS/ML

Tools

GPUs
Distributed training

Job description

AMD is hiring an AI Research Scientist specializing in reinforcement learning for post-training and interactive learning of large generative models applied to engineering and hardware tasks. You will design RL algorithms, run studies, and collaborate with infra and product teams to land methods that improve task success and stability.

The role requires a PhD (preferred) and a strong publication record, with hands-on experience training RL or preference-based models at scale, especially with LLMs

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RL Research Scientist for Post-Training LLMs & Code Models
RL Research Scientist for Post-Training LLMs & Code Models

Advanced Micro Devices • Santa Clara (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
AI Research Scientist, Reinforcement Learning (LLM) and Post-Training
AI Research Scientist, Reinforcement Learning (LLM) and Post-Training

AMD • Santa Clara (CA)

On-site
USD 150,000 - 230,000
Benefits at a glance
AI Research Scientist, Reinforcement Learning (LLM) and Post-Training
AI Research Scientist, Reinforcement Learning (LLM) and Post-Training

Advanced Micro Devices • Santa Clara (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Research Engineer: RL & Post-Training LLM Systems
Research Engineer: RL & Post-Training LLM Systems

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
ML Engineer — LLM Post-Training & RL Specialist
ML Engineer — LLM Post-Training & RL Specialist

NewsBreak • Mountain View (CA)

On-site
USD 130,000 - 160,000
Health, dental, and vision care
401(k) plan with company matching
Paid time off and holidays
LLM Post-Training Research Scientist (SFT & RLHF)
LLM Post-Training Research Scientist (SFT & RLHF)

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Health, dental & vision coverage
Retirement benefits
Learning and development stipend
+2
Generative AI Research Scientist: LLM Post-Training
Generative AI Research Scientist: LLM Post-Training

Scale AI, Inc. • New York (NY), Northern (KY)

Hybrid
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
RL & Inference ML Systems Engineer for Engineering AI
RL & Inference ML Systems Engineer for Engineering AI

AMD • Santa Clara (CA)

On-site
USD 180,000 - 250,000
AMD benefits
AI Research Engineer
AI Research Engineer

Alexander Chapman • San Francisco (CA)

On-site
USD 150,000 - 230,000
Lead RL Infra Engineer - Scalable GPU Training Platforms
Lead RL Infra Engineer - Scalable GPU Training Platforms

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 120,000 - 170,000