Position Overview
We’re looking for a talented Staff Software Engineer to join our Agentic Platform team—the group responsible for building the company‑wide infrastructure that enables every engineering team to safely build AI‑powered features. This is a rare opportunity to shape foundational AI/LLM platform capabilities from the ground up at a company that’s deploying real‑world AI agents today.
The Agentic Platform provides shared primitives, hosted agent execution, and operational tooling so any team can build AI‑powered workflows—from simple summarization to complex multi‑turn conversational agents. Our vision is to enable any engineer to build production‑ready AI features without becoming an AI expert.
The Platform handles the hard infrastructure problems—provider abstraction, safety guardrails, observability, prompt lifecycle management, and evaluation systems—so Product Teams can focus on their domain logic. Reporting to the Director of Engineering, you’ll partner closely with the Principal Engineer leading platform architecture while collaborating with product teams across the company who will consume your platform.
Who You Are
You’re a platform engineer at heart—an expert who builds invisible infrastructure that handles enormous complexity. You have a track record of designing systems that other engineers love to use, balancing powerful abstractions with practical simplicity.
You understand the unique challenges of LLM systems: non‑deterministic outputs, rapidly evolving provider SDKs, safety requirements, and systematic quality measurement. You’re excited to tackle churn containment—building stable APIs that absorb monthly model releases and quarterly SDK updates.
The ideal candidate is a hands‑on, outcome‑oriented engineer with extensive experience building platform infrastructure, developer tools, or distributed systems, and with a strong partnership with internal customers.
What You Will Do
- Design and build core platform primitives including provider abstraction layers (OpenAI, Anthropic, Google), structured output validation, streaming infrastructure, and token management systems.
- Own safety and compliance infrastructure including composable guardrail systems, PII detection/redaction, audit logging, and privacy‑first observability.
- Build evaluation infrastructure to enable systematic quality measurement for non‑deterministic LLM outputs—datasets, scorers (exact match, LLM‑as‑judge, schema validation), CI/CD integration, and regression detection.
- Lead churn containment strategy—design provider adapters and SDK architecture that absorb rapidly‑changing LLM provider SDKs without breaking consuming applications.
- Architect prompt lifecycle management systems including version control, Langfuse integration, GitHub‑based review workflows, and deployment pipelines.
- Design Agent‑as‑a‑Service infrastructure for long‑running async tasks using AWS EventBridge, DynamoDB, and PostgreSQL.
- Collaborate with consuming teams to understand their needs, onboard them to the platform, and provide technical support.
- Influence architecture, technology selections, and engineering standards across the broader organization.
- Create reference implementations and technical documentation that enable other engineers to adopt the platform.
- Champion quality engineering practices including comprehensive testing, type safety, and observability.
Required Skills & Competencies
- 8+ years of software engineering experience with significant time spent building platform infrastructure, developer tools, SDKs, or distributed systems.
- Production experience with LLM/AI systems—building and operating systems using OpenAI, Anthropic, or similar providers.
- Strong TypeScript expertise; design APIs consumed by other TypeScript developers.
- Experience designing APIs and abstractions that other engineers love to use.
- Understanding of safety and compliance in AI systems—PII handling, guardrails, audit logging, responsible AI practices.
- Experience with event‑driven architectures and async processing patterns (EventBridge, SQS, or similar).
- Understanding of observability and monitoring for distributed systems—metrics, tracing, alerting, debugging production issues.
- Strong communication and technical writing skills; ability to document systems clearly and work with internal customers across multiple teams.
- Track record of technical leadership without formal management—focusing on architecture, mentoring engineers, driving technical decisions.
- Experience with cloud infrastructure (AWS preferred: Fargate, DynamoDB, RDS, S3, EventBridge).
Preferred
- Experience building SDK or platform products consumed by multiple teams.
- Experience with prompt engineering, prompt management systems, or LLM evaluation frameworks.
- Familiarity with NestJS, Prisma, or similar TypeScript backend frameworks.
- Experience with streaming architectures (SSE, WebSockets) for real‑time AI applications.
- Background in building multi‑tenant platform infrastructure.
- Experience with hexagonal architecture / ports and adapters patterns.
- Contributions to open‑source LLM tooling or frameworks.
Technical Environment
- Languages: TypeScript (primary)
- Frameworks: NestJS, OpenAI Agents SDK, Vercel AI SDK
- Databases: PostgreSQL (Prisma ORM), DynamoDB, Redis
- Infrastructure: AWS (Fargate, EventBridge, S3, Parameter Store), Docker
- Observability: Langfuse, NewRelic, Coval
- Testing: Vitest
- CI/CD: GitHub Actions, SonarQube
- LLM Providers: OpenAI, Anthropic (with architecture for additional providers)
- Coding Agents: Claude, Codex, Gemini
Compensation
- Base Salary: $160,000 to $190,000 + 10% Bonus
- Benefits:
- 401(k) plus match
- Dental insurance
- Health insurance
- Vision insurance
- Paid Time Off
EEO Statement
All your information will be kept confidential according to EEO guidelines. A Place for Mom uses E-Verify to confirm the employment eligibility of all newly hired employees. To learn more about E-Verify, including your rights and responsibilities, please visit www.dhs.gov/E-Verify.