An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Rox in San Francisco is seeking a researcher to lead applied AI initiatives in enterprise settings. You’ll design and run research programs tied to four priority problems, build evaluation frameworks, and work on memory, retrieval, and context systems with elite engineers.
You’ll translate findings into production infrastructure, helping define Rox Research’s next steps. A strong research instinct and ability to ship are essential; a PhD is not required.
Foundation models are commoditizing. Defensibility comes from specialized models, proprietary training signals, and evaluation ownership. Every applied AI company we benchmark against like Decagon, Harvey, Sierra, Cursor has already moved. The window to claim frontier applied AI for revenue is closing in the next few months.
Rox is in market. We run agents against enterprise data at scale, every day. We see exactly where research meets production and where the data is dirty, state is changing, and being wrong costs (a lot of) money.
The Applied Research team exists to close that gap permanently.
These are not benchmark problems. They have real SLAs and real customers depending on them.
You have spent real time thinking about how agents fail in practice, not just on benchmarks. You have built evaluation systems and know exactly where standard approaches break down. You can write code well enough to implement your own ideas, run your own experiments, and ship things that make it into production.
You move fast. The environment changes monthly and the team ships continuously.
Particularly relevant: agent evaluation and behavioral benchmarking; retrieval-augmented generation and knowledge graph systems; RL applied to real-world agent behavior; production ML systems (latency, reliability, observability); post-training and model adaptation for production use cases.
A PhD is not required. Strong research instincts and the ability to ship are.
First few weeks: you understand Rox's architecture, where the production problems are, and where the research gaps are. You have opinions and you share them.
First few months: you are running experiments that directly inform how we build. Something you worked on is in production.
Over time: you are defining the research agenda for the most interesting applied AI problem in the enterprise. The systems you build are things no one else has built before, because no one else has the structural data position to build them.
We are at an unusual moment. Large enough to have real scale, real customers, and genuinely interesting research problems. Small enough that you are one of a handful of people shaping what the Applied Research function looks like and what it prioritizes.
The team is extraordinary: IMO, IOI, and ICPC medalists, researchers from DeepMind and OpenAI. The feedback loop is a live enterprise system, not a leaderboard. If that's not more interesting to you than publishing for the sake of publishing, this probably isn't the right fit.
San Francisco, onsite. We relocate exceptional people.