AI Researcher – RL & On-Device Models

Rnb Consultancy

San Francisco (CA)

On-site

USD 200,000 - 250,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Relocation support
Immigration support

Job summary

San Francisco AI Lab is hiring to advance on-device agents that run privately on consumer devices. You will own end-to-end model training, from pretraining decisions to distillation and quantization, with a focus on small, efficient models that outperform their size.

You'll work across data mix, objectives, and production runtime, with a clear metric for your stream and checkpoints that reach production. Relocation and immigration support are provided.

Qualifications

  • Hands-on depth across the training stack: pretraining, SFT/RLHF/DPO/GRPO and quantization.
  • Experience with RL, on-device deployment, or both; systems fluency is a plus.
  • Design experiments around metrics that matter and ship checkpoints.
  • Experience in frontier labs or startups during high growth.

Responsibilities

  • Lead the training of a model family powering on-device agents, including post-training and distillation.
  • Own end-to-end capabilities: data mix → objective → evals → production runtime.
  • Manage your own workstream with measurable checkpoints sent to production.
  • Run experiments and publish actionable write-ups to inform the team.
  • Collaborate closely with infra, product, partnerships, and researchers.

Skills

Training stack depth
On-device deployment
Systems fluency (memory, GPU)
Experiment design
Startup/hypergrowth experience

Job description

Done training models that disappear behind an API? Put yours on phones, laptops and glasses instead.

A San Francisco AI lab is building trustworthy, consumer-grade agents that run privately on the device. With their small action models and execution layer, an agent can read what's on screen and operate any app the way a person would, without APIs or custom integrations per app. They've held the top spot on the industry's leading mobile-agent benchmark since late 2025, work with a leading mobile chipmaker and a global device maker on edge deployments, and are backed by top-tier VCs and senior leaders from frontier labs. The founder has built and funded an agent startup before. The team has more than doubled in the past half year, and a new round is expected to close before the end of the year. Two lanes: consumer apps and partnerships with device makers.

The core challenge: squeezing frontier-level capability into the tight compute and memory limits of a phone or laptop.

What you'll own
  • Leading the training of a model family that powers the agents: pretraining recipe decisions, post-training (SFT, RLHF, DPO, GRPO and beyond), distillation, quantization, and every trick that helps a small model outperform its size
  • One or more capabilities end to end: data mix → objective → evals → shipped into a production on-device runtime
  • Your own workstream, measured on one clear metric, ending in a checkpoint that goes to production
  • Experiments and write-ups the whole team builds on, and the discipline to drop ideas that don't move the needle
  • Close work with infra and product engineers, the partnerships team and fellow researchers
Your first 90 days
  • Day 30: you've reproduced a recent training run end to end and named the three bets with the most leverage
  • Day 60: you're leading a workstream and have shipped a checkpoint that outperforms the best one so far
  • Day 90: what you built is running in a partner's build
What you bring
  • Hands-on depth across the training stack: pretraining, SFT/RLHF/DPO/GRPO, distillation and quantization
  • Tricks for small models: distilling from frontier teachers, MoE at small scale, KV-cache compression, speculative decoding
  • Experience with RL, on-device deployment, or both; systems fluency (memory, GPU optimization) is a big plus
  • You design experiments around metrics that matter, and you ship checkpoints, not just papers
  • You've lived through hypergrowth at a frontier lab, or you were a founder or one of the first hires at a startup that took off
  • Relevance over pedigree: exceptional work counts more than a big-name logo
Bonus points
  • Published work on RL, on-device ML or efficient inference
  • Models shipped to production on-device runtimes against hardware release dates
  • ML at consumer scale
What's in it for you
  • $200k–$250k base + equity
  • Your model running on real consumer devices
  • Relocation and immigration support
Good to know
  • Full-time, in person in San Francisco, 9-9-6
  • Process: intro call → short talk with the founder → ~60-min technical conversation about design and first principles (no live coding, no puzzles) → half a day onsite with the team → a work trial in person on an actual problem → offer
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Researcher — Browser Agents & Product Impact
AI Researcher — Browser Agents & Product Impact

MULTI·ON • San Francisco (CA)

On-site
USD 100,000 - 130,000
Competitive compensation package
Significant equity
Access to substantial compute resources
AI Researcher
AI Researcher

AGI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 280,000
Research Engineer – Evals
Research Engineer – Evals

Rnb Consultancy • San Francisco (CA)

On-site
USD 200,000 - 250,000
Equity
Relocation support
Visa/immigration assistance
Applied AI Engineer – Long-Horizon Agents
Applied AI Engineer – Long-Horizon Agents

Rnb Consultancy • San Francisco (CA)

On-site
USD 250,000 - 500,000
Equity 1–5%
Visa sponsorship (O-1)
Founding-level ownership
+2
Research Scientist – RL Post-Training for Agents
Research Scientist – RL Post-Training for Agents

Rnb Consultancy • San Francisco (CA)

On-site
USD 250,000 - 500,000
Visa sponsorship
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Applied Research - RL & Agents
Applied Research - RL & Agents

Prime Intellect, Inc. • San Francisco (CA)

On-site
USD 150,000 - 300,000
Salary + equity incentives
Flexible work—SF or hybrid-remote
Visa sponsorship & relocation support
+2
Applied Research - Forward-Deployed
Applied Research - Forward-Deployed

Prime Intellect, Inc. • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Cash compensation range: $150-300k
Flexible work (San Francisco or hybrid
Visa sponsorship & relocation
+2
AI Product FDE
AI Product FDE

AGI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Relocation support
Immigration assistance
In-person in SF
AI Product FDE
AI Product FDE

Maven Ventures • San Francisco (CA)

On-site
USD 160,000 - 230,000