- You’ll join a genuine 0→1 team on the ground floor of one of the company’s biggest new bets. This seat is specifically product-focused: you’ll own agent quality, ship agents that do real audit work, and work alongside practitioners
- Depending on your experience and what you’re looking to own, you may join building within a major agent area, owning one end-to-end, or setting technical direction for agentic audit work across the team
- We’re hiring across all levels and will calibrate during interviews based on scope and demonstrated experience
- Make agent judgment repeatable: run error analysis on real testing data and turn findings into concrete fixes
- Tradeoffs such as quality/latency/cost across a long multi-phase run
- Build structured-output pipelines that turn model output into real audit artifacts
- Take ambiguous problem statements and turn them into a plan, a shipped feature, and a clear read on what was cut and why
- Work directly with an embedded subject matter expert and with design-partner firms, turning their feedback into agent changes within days
- Expand agent coverage into new controls and new areas of internal audit
- Product-minded and full-stack: you’ve shipped LLM-backed features to production against real users, and you measure yourself on whether they got used
- You’re fluent in evals and error analysis, and you apply them in service of shipping something practitioners trust
- You have real opinions on model selection, prompting, and orchestration tradeoffs, and you can defend them with evidence rather than vibes
- Energized by 0→1 work: you’d rather define the problem than inherit a spec, and you don’t stall on ambiguity
- Strong instincts for human-in-the-loop design
- A genuine team player across the organization, not just within engineering: you’ll work daily with PM, design, domain experts, and customer-facing teams, and you treat that as the best part of the job
- Ship fast without leaving a mess: your code is reviewable, tested where it counts, and instrumented
- Able to internalize a hard domain fast. You don’t need to know SOX today, but you’ll understand it well enough to make the right product calls
- Own a major agent area end-to-end, from how the agent reasons about a class of controls through to the artifact a reviewer signs
- Set the evals and error-analysis practice for the team’s agent work, and decide what evidence justifies shipping a change or rolling it back
- Collaborate with PMs and designers to shape roadmaps and define architectural tradeoffs, including where the agent acts and where the auditor decides
- Own the harder model and orchestration judgment calls across a long multi-phase run
- Mentor other engineers and raise the bar on 0→1 execution and applied eval rigor
- Drive agent initiatives that reach beyond Internal Audit and influence how agents are built across Fieldguide
- Set and champion engineering standards for agent reliability, reproducibility, and defensibility
- Partner with engineering and product leadership to define long-term technical strategy for agentic audit work
- Serve as a trusted advisor to leaders across Engineering, Product, and Design
- Represent Fieldguide externally through writing, speaking, and open-source contributions
Benefits
Applied AI skillset: evals, error analysis, and model-selection decisions you owned and can explainShipped LLM-backed product features to production against real usersExperience working directly with customers, and comfort being in the room when they use what you builtComfortable full-stack, with enough backend depth to work in agent orchestrationA collaborative mode that works across PM, design, and domain expertsAutonomy working from an ambiguous specStructured-output work including schema contracts, generating real artifacts from model outputPython, TypeScript, React, Postgres, Hasura, GraphQLStartup experience, as a founder or as an early engineerHands-on eval experience (Langfuse, Braintrust, LangSmith, Arize Phoenix, or comparable)Temporal or comparable durable-execution / workflow orchestrationBackground in internal audit, SOX, accounting, or another regulated domainDocument processing, including PDF and Excel manipulation and annotationA 0→1 track record: things you started where no scaffolding existed