Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Zep AI in San Francisco seeks an engineer who builds agents and trains the models they run on. You will own the memory extraction and retrieval loop, finetune models, and push systems to production. You’ll report to the founder and join a team with experience at Scale AI, Dropbox, and McKinsey.
You will work in a distributed team, write production-grade Python, design experiments, and build eval harnesses to catch regressions before release. A Master’s in CS or equivalent is required.
Zep manages, governs, and serves agent memory at enterprise scale. Enterprises build on Zep to run reliable, personalized agents across the business: millions of Context Graphs, served in under 200ms, inside their own VPCs and cloud deployments. Customers include Samsung, Zscaler, Twin Health, HoneyBook, and NASDAQ 100 and Fortune 500 technology companies. We also build Graphiti, our open-source context graph framework (30K+ GitHub stars).
What our customers' agents can reason about depends on the memory we retrieve and the memory we write. You own that loop. You build agents that improve retrieval. You finetune the models that extract memory, and the models that power those agents. You run the experiments and ship the result as production code.
We're hiring an engineer who builds agents and trains the models they run on.
We will measure you on whether retrieval and memory quality move, and on what reaches production.
You’ll report to our founder, Daniel (2x founder, engineer, former head of ML at SparkPost), and join a team with pedigree at Scale AI, Dropbox, ActiveCampaign, DroneDeploy, and McKinsey.
We're a small, distributed team that works closely together. We pair on hard problems, review each other's designs, and treat learning as part of the job rather than something that happens after hours. We ask a lot of questions: of customers, of teammates, of our own assumptions. When we find pain, we go fix it.
We expect the same back: ask questions early, push back when you disagree, and care about the people on the other end of the API.
In your first week you pick up an open retrieval or memory-extraction problem and frame it as an experiment. By day 30 an agent or a finetune you built is in the product loop, or you have ruled the approach out and written up why. By day 90 you are three cycles in and the eval harness you built runs on every release.