Research Intern, Agent RL Training

GoTo Meeting

Mountain View (CA)

On-site

USD 48,216 - 68,880

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NewsBreak in Mountain View, CA is hiring a Research Intern for their Agent RL Training team. The role involves collaboration with a mentor to drive research on applying large language models to core business operations. Responsibilities include running SFT experiments, curating training datasets, and contributing to public publications.

Ideal candidates are research-driven individuals with strong Python and PyTorch skills. The internship offers an hourly pay between $35 and $50.

Qualifications

  • Willing to put in extra hours to push projects across the finish line.
  • Genuine passion for research, reads papers, and tinkers with models.
  • Independently capable of end-to-end model SFT.

Responsibilities

  • Collaborate with mentor to identify research directions for applying LLMs.
  • Run end-to-end SFT experiments on LLM-based agents.
  • Contribute to public publications during your internship.

Skills

Python
PyTorch
Research motivation
End-to-end model SFT

Tools

CUDA
Triton
OpenRLHF

Job description

About NewsBreak

Founded in 2015, NewsBreak is the Content Intelligence platform shaping the future content economy. With over 40 million monthly active users, our flagship platform delivers highly personalized local news and information powered by advanced AI, recommendation systems, and adtech.

Recognized by Fast Company as #32 on the Top Workplaces for Innovators, we’re proud to be Great Place to Work certified and home to a dynamic team of technologists, product innovators, and business leaders who are passionate about solving meaningful challenges at scale.

Together, we reached unicorn status in 2021, and we remain committed to continuing this high-growth trajectory with the right team to fulfill our mission: building the infrastructure layer for content intelligence.

About the Role

We are looking for a Research Intern to join our Agent RL Training team. You will be paired with a full-time employee as your mentor, working together to explore, from zero to one, how to apply large language models to NewsBreak’s core business, including content understanding, recommendation, agentic web browsing, and autonomous multi-step task completion.

This is a hands-on research role. You are expected to independently drive experiments, propose novel ideas, and iterate quickly. We value self-starters with deep intellectual curiosity and the drive to push boundaries in LLM post-training and agent capabilities.

Location: Onsite in Mountain View, CA office

What You’ll Work On
  • Collaborate with your full-time mentor to identify high-impact research directions for applying LLMs to NewsBreak’s products
  • Independently run end-to-end SFT experiments on LLM-based agents, and assist with RL-related exploration such as reward design and training iteration
  • Curate and build high-quality training datasets: instruction-following, preference pairs, agent trajectories, and synthetic data
  • Contribute to public publications; we encourage and support top-venue submissions during your internship
Requirements
  • Highly motivated and committed: willing to put in extra hours when needed to push projects across the finish line
  • Genuine passion for research: you read papers for fun, tinker with models on weekends, and care deeply about advancing the field
  • Independently capable of end-to-end model SFT: with basic understanding of RL-based post-training methods (RLHF, DPO, PPO, GRPO, etc.)
  • Excellent taste in model behavior: able to reason about what "good" looks like across user-facing domains and articulate why
  • Strong Python and PyTorch skills
Preferred Qualifications
  • Publication at a top-tier venue (NeurIPS, ICML, ICLR, ACL, EMNLP, or equivalent)
  • Experience with multi-node distributed training (FSDP, DeepSpeed, Megatron-LM)
  • Proficiency in writing custom GPU kernels with Triton or CUDA
  • Experience building synthetic data pipelines for agent training
  • Familiarity with open-source RL frameworks: TRL, OpenRLHF, veRL/vLLM

Hourly Pay: $35- $50

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Intern, Agent RL Training
Research Intern, Agent RL Training

NewsBreak • Mountain View (CA)

On-site
Research Intern — LLM Agents & RL Training
Research Intern — LLM Agents & RL Training

NewsBreak • Mountain View (CA)

On-site
LLM Research Intern: Agent RL Training & Innovation
LLM Research Intern: Agent RL Training & Innovation

GoTo Meeting • Mountain View (CA)

On-site
Research Intern
Research Intern

Quadrillion Labs • New York (NY)

On-site
USD 300,000 - 500,000
Medical, dental, and vision insurance
Lunch and dinner covered
Other varied stipends
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Competitive cash and equity compensation (>90th percentile)
Ownership and autonomy
Health, vision, dental benefits
+4
Machine Learning Engineer, LLM Post-Training
Machine Learning Engineer, LLM Post-Training

NewsBreak • Mountain View (CA)

On-site
USD 130,000 - 160,000
Health, dental, and vision care
401(k) plan with company matching
Paid time off and holidays
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Machine Learning Researcher
Machine Learning Researcher

Multicoin • San Francisco (CA)

On-site
USD 250,000 - 350,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Research Intern RL & Post-Training Systems, Turbo (Fall 2026)
Research Intern RL & Post-Training Systems, Turbo (Fall 2026)

Togetherai • San Francisco (CA)

Hybrid
USD 79,900 - 86,788
Competitive compensation
Housing stipends
Other competitive benefits
RESEARCHER, POST-TRAINING
RESEARCHER, POST-TRAINING

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000