Get more replies from employers
Send a job-specific resume in minutes.
Varick Agents LTD. is seeking a production-focused ML engineer to own post-training loops for enterprise-grade AI agents in our San Francisco environment. You will convert deployment data into models, evaluation data, and iterative improvements to boost agent accuracy beyond demo-grade levels.
The role emphasizes applied ML, real-world usage, and cost-aware inference, with a fast-paced small-team culture and opportunities to influence how agents operate inside major enterprises.
Varick builds AI agents that take over real operational workflows inside the world's largest enterprises. Our forward-deployed engineers and strategists embed with client teams, map how work actually runs, then our engineers build the agents into production. We're venture-backed, revenue-generating from day one, and already in production inside several of these companies. You'll join an elite team from Meta, AWS Bedrock, Citadel Securities, McKinsey, BCG, Stanford, and more.
You make our agents better at the actual work. You own the post-training loop: taking real workflow data from our deployments and turning it into models and evals that push agent accuracy where a demo-grade model would fail on the tenth run. This is applied, production-facing ML, not research for its own sake.
Own fine-tuning, preference optimization, and post-training pipelines that lift agent performance on real enterprise workflows
Build the eval suites and golden sets that measure whether a change actually helped, before it ships to a client
Turn traces and failure cases from production into training data and targeted improvements
Work with the platform team on the model layer: routing, distillation, cost and latency budgets
Keep a clear-eyed view of where fine-tuning earns its keep versus better prompting, retrieval, or scaffolding
Strong applied ML background with hands-on post-training experience (SFT, RLHF/DPO, or similar) on LLMs
At least one model or system you took to real users, and a clear account of how it behaved in production
Fluency with the modern training stack and a working grasp of inference economics
Evaluation discipline: you trust measured results over vibes
Comfort moving fast in a small team, with judgment on where rigor matters
Published research or open-source work in LLMs and agents
Experience with agentic evals, tool-use training, or long-context work
A track record of taking research into production
Meaningful equity, real ownership, direct access to how the world's largest enterprises actually run, flexible PTO, free lunch and dinner in the office, Ubers home if you're staying late, and monthly team dinners.
On-site in our San Francisco Financial District office, 5-6 days a week. Comfortable with startup hours, for startup upside.
Equal opportunity employer. Employment is subject to a standard confidentiality and non-disclosure agreement.