Stand out for this role — generate a tailored resume and cover letter in about a minute.
Pulley is seeking a senior AI engineer to build the intelligence behind our permitting platform. You’ll own AI-powered features from user interaction to production, turning messy inputs into structured outputs and building robust evals.
You will work with agents and large language models in a production setting, shaping product decisions and engineering standards. You’ll implement retrieval-augmented workflows, document understanding, and agentic processes, guiding data quality, latency, and
You’re product-minded: you care whether the thing you built actually solved the customer’s problem, and you’ll talk to users to find outYou default to ownership: when something is broken or missing, your instinct is to fix it, not to file it as someone else’s problemYou have strong opinions about quality and velocity and don’t treat them as a tradeoff—you look for the tools, abstractions, and processes that buy bothYou thrive in ambiguity—you’d rather define the right problem than execute a spec, and you’re energized rather than paralyzed when the path isn’t laid outYou’re rigorous about what “working” means—you don’t trust a demo, you trust an eval, and you build the measurement before you build the featureTrack record of owning an LLM-powered product surface end-to-end: requirements through production, including the unglamorous parts—data quality, eval design, cost and latency, failure handling4+ years of software engineering experience, with a substantial portion building production LLM or ML systemsAbility to architect durable systems while making pragmatic tradeoffsReal experience building with AI coding agents—not just autocomplete; you’ve shipped work where agents did substantial implementation under your directionDeep hands-on experience with large language models in production—prompting, retrieval-augmented generation, structured extraction, tool use and agentic workflows, and knowing when each is the wrong toolBased in the San Francisco Bay Area and willing to work in person 4 days a weekExperience designing evals and otherwise making LLM-powered features reliable in productionExperience with document understanding at scale—OCR, layout-aware parsing, or vision-language models over scanned PDFs, drawings, or formsExperience fine-tuning models or building data pipelines to produce training and eval sets from real-world usageExperience in construction tech, govtech, proptech, or another domain where the hard part is messy real-world documents and processesExperience with modern full-stack development—we use TypeScript, React, and Google Cloud—and an appetite for working in the application code that puts AI features in front of usersStartup experience at the stage where you helped build the team, not just the productExperience mentoring engineers or leading technical direction across teams