Done training models that disappear behind an API? Put yours on phones, laptops and glasses instead.
A San Francisco AI lab is building trustworthy, consumer-grade agents that run privately on the device. With their small action models and execution layer, an agent can read what's on screen and operate any app the way a person would, without APIs or custom integrations per app. They've held the top spot on the industry's leading mobile-agent benchmark since late 2025, work with a leading mobile chipmaker and a global device maker on edge deployments, and are backed by top-tier VCs and senior leaders from frontier labs. The founder has built and funded an agent startup before. The team has more than doubled in the past half year, and a new round is expected to close before the end of the year. Two lanes: consumer apps and partnerships with device makers.
The core challenge: squeezing frontier-level capability into the tight compute and memory limits of a phone or laptop.
What you'll own
- Leading the training of a model family that powers the agents: pretraining recipe decisions, post-training (SFT, RLHF, DPO, GRPO and beyond), distillation, quantization, and every trick that helps a small model outperform its size
- One or more capabilities end to end: data mix → objective → evals → shipped into a production on-device runtime
- Your own workstream, measured on one clear metric, ending in a checkpoint that goes to production
- Experiments and write-ups the whole team builds on, and the discipline to drop ideas that don't move the needle
- Close work with infra and product engineers, the partnerships team and fellow researchers
Your first 90 days
- Day 30: you've reproduced a recent training run end to end and named the three bets with the most leverage
- Day 60: you're leading a workstream and have shipped a checkpoint that outperforms the best one so far
- Day 90: what you built is running in a partner's build
What you bring
- Hands-on depth across the training stack: pretraining, SFT/RLHF/DPO/GRPO, distillation and quantization
- Tricks for small models: distilling from frontier teachers, MoE at small scale, KV-cache compression, speculative decoding
- Experience with RL, on-device deployment, or both; systems fluency (memory, GPU optimization) is a big plus
- You design experiments around metrics that matter, and you ship checkpoints, not just papers
- You've lived through hypergrowth at a frontier lab, or you were a founder or one of the first hires at a startup that took off
- Relevance over pedigree: exceptional work counts more than a big-name logo
Bonus points
- Published work on RL, on-device ML or efficient inference
- Models shipped to production on-device runtimes against hardware release dates
- ML at consumer scale
What's in it for you
- $200k–$250k base + equity
- Your model running on real consumer devices
- Relocation and immigration support
Good to know
- Full-time, in person in San Francisco, 9-9-6
- Process: intro call → short talk with the founder → ~60-min technical conversation about design and first principles (no live coding, no puzzles) → half a day onsite with the team → a work trial in person on an actual problem → offer