Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Ema in Vancouver is building AI employees that carry out complex workflows across enterprise applications. This role focuses on designing context, tools, and orchestration for multi-step agents, spanning documents, slides, images, audio and video, with careful test decisions on latency and cost.
You’ll join researchers and engineers across the Bay Area, Vancouver and India, contributing to post-training methods, evaluation, and production deployment.
A master’s or PhD in a relevant field, or equivalent work or research experience. Papers, substantial open-source contributions, trained models and well-documented experiments can demonstrate that depthStatistical judgment. You can size an experiment, choose meaningful baselines and held-out tests, and account for variation across tasks, seeds and repeated runs. You can distinguish a real improvement from judge bias, data leakage or a benchmark shortcutEvidence of zero-to-one ownership. A system, model or research project you took from an ambiguous problem to a working result. We want to understand your contribution, the tradeoffs you made, and what changed when the work met real users or realistic tasksThese are project-specific strengths; post-training experience is optional, and no candidate needs the entire list. Bring a repository, paper, model or technical write-up that lets us examine how you think and what you builtDepth you can defend. Substantial work in at least one of agent/tool-use systems, post-training, reward modeling or RL environments, retrieval and memory, or evaluation design. Be ready to explain the mechanism, the alternatives you rejected and the failure modes you found. One area you can teach us beats five you’ve touchedProduction engineering judgment. You can debug across the model and system boundary, isolate a failure, and turn the result into maintainable production codeHonest measurement. You would rather retire your own approach after a clean negative result than ship an improvement that disappears under a stronger evaluationInteractive agent environments/harnesses for software engineering, web or tool use; large-scale trace analysis, data curation or synthetic generationPractical security work on prompt injection, data governance or permission boundaries for agents that can act and improve themselvesOpen-model post-training with TRL, veRL, OpenRLHF, or similar or a custom loop, especially debugging reward hacking or unstable optimizationDesigning systems to support complex, long-horizon agent work across a multitude of modalities and platformsServing with vLLM or SGLang, distillation, quantization, or multi-node GPU training