Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Armadin, based in Palo Alto, is seeking a senior researcher to own the post-training strategy and to translate security research into scalable model capabilities. You will drive fine-tuning, reward modeling, and efficiency benchmarks, ensuring reliability and impact across production environments.
The role requires a PhD or equivalent experience with hands-on SFT, RLHF/RLAIF, and distillation. You will collaborate across teams to push frontier research into practical, enterprise-grade solutions.
At Armadin, we’re a group of engineers, researchers, and hackers on a mission to redefine what proactive security can do in the AI era. Cyberattacks are becoming autonomous and relentless and we believe defending against threats before they materialize is one of the most powerful ways to protect the institutions the world depends on.
We’re building autonomous proactive security from the ground up, reinforcing the tradecraft of elite red teamers into purpose-built security models and agents that discover risk and remediate it before organizations are breached.
Led by Kevin Mandia, founder of Mandiant ($5.4B exit to Google), our team brings together researchers and engineers from Google, xAI, Meta, Stanford, and MIT to reinvent security for an adversary that never sleeps.
Own Post-Training Strategy: Drive the post-training roadmap (fine-tuning, preference optimization, reward modeling, RL, distillation) to make models more capable, reliable, and aligned.
Push Efficiency & Evals: Make models serve reliably at scale, and build the benchmarks that measure quality and catch regressions.
Ground Research in Reality: Turn our security experts’ tradecraft into model capabilities and ship your techniques into production.
Stay at the Frontier: Track post-training and efficiency research and bring the best ideas in.
Research Pedigree: PhD in CS, ML, or related field, or equivalent experience with a strong track record.
Post-Training Depth: Hands‑on experience with SFT, RLHF/RLAIF, DPO, and reward modeling.
Practical Results: You train, fine‑tune, and evaluate models, and make research work in practice.
Engineering Strength: You implement your own ideas and run them at scale.
Range & Judgment: You drive strategy broadly or go deep on one hard problem as needed.
Collaboration: You work well across teams and care about the outcome, not just your slice.