San Francisco Hybrid (2 days a week)
The Company
A high-growth, venture-backed SaaS company is building AI-powered products for complex, knowledge-intensive workflows. The business has established market traction and is making a significant investment in the infrastructure required to deliver reliable applied AI at scale.
The Opportunity
This is a principal-level opportunity to help define how production AI systems are designed, evaluated, monitored, and scaled across the company. You will have broad technical ownership and direct influence over a platform that is central to the next phase of the product roadmap.
The Role
You will architect and build the infrastructure behind production LLM applications, with a focus on evaluation, ingestion, observability, retrieval, and operational reliability. You will combine hands-on engineering with technical leadership, setting standards and helping multidisciplinary teams move from prototypes to durable customer-facing systems.
What You’ll Do
- Architect and scale the company’s core AI platform and supporting infrastructure.
- Build evaluation frameworks, monitoring systems, ingestion pipelines, and observability tooling for production AI applications.
- Design and improve retrieval, embedding, prompting, and model-interaction workflows.
- Lead AI projects from early experimentation through deployment, measurement, and ongoing optimization.
- Establish engineering standards for testing, versioning, quality control, and system reliability.
- Partner with product, design, engineering, and domain specialists to translate complex workflows into effective AI products.
- Provide technical direction and raise the quality of AI engineering across the wider organization.
What You’ll Bring
- Deep software engineering experience across backend systems, infrastructure, distributed systems, or platform engineering.
- A track record of shipping AI or machine learning systems into production.
- Strong Python skills and the ability to perform well in rigorous coding and system design discussions.
- Experience with areas such as LLMs, retrieval systems, agents, NLP, evaluations, observability, or data ingestion.
- The judgement to balance experimentation, delivery speed, technical quality, and long-term maintainability.
- A high-ownership approach and the ability to lead technically without relying on formal management authority.
- Clear communication skills and confidence working across engineering, product, design, and business teams.
What This Role Requires
- Production experience building or operating AI-enabled software systems.
- Strong backend or platform engineering fundamentals.
- Ability to design and own complex systems from initial architecture through production operation.
- Willingness to work on-site in San Francisco.
Why Join
- Own a strategically important AI platform with company-wide product impact.
- Solve complex production challenges across evaluation, reliability, retrieval, data, and infrastructure.
- Join during a major scaling phase with strong compensation, equity, and meaningful technical influence.