We are looking for an Architect / Principal Software Engineer to lead the technical direction of our AI Engineering pods. You will split your time between hands‑on engineering (about 30% to 40% prototyping core workflows, building reference implementations, and reviewing code) and higher‑level system architecture, technical standards, and mentorship across teams. You will work closely with US‑based Product, Security, Data, and Infrastructure teams to make sure our agentic systems are reliable, fast, and secure. Your main focus will be designing predictable multi‑agent workflows, low‑latency LLM orchestration services, and retrieval pipelines that run cleanly in production.
Req.#1098579355
Responsibilities
- Set architectural direction across multiple AI pods, making clear trade‑offs between model capabilities, cost, latency, and operational complexity
- Spend roughly a third of your time in code: writing backend AI services, building orchestration patterns, and defining integration contracts with frontend teams
- Design agent workflows with streaming responses, predictable error handling, and low latency for chat and automated tasks
- Establish shared standards for Model Context Protocol (MCP) servers, tool schemas, and context management
- Work with Okta Security and Infrastructure to enforce least‑privilege tool execution, tenant isolation, and secret handling within agent loops
- Define observability requirements and service‑level objectives, including tracing with tools like LangSmith or Arize Phoenix
- Mentor Staff and Senior engineers through direct design reviews and technical guidance
Requirements
- 10+ years of software engineering experience, with a background in designing, scaling, and operating distributed backend systems
- 6+ years of Python experience, especially with FastAPI, AsyncIO, and modern typing practices. Python is our primary language for backend and AI services
- 3+ years building production systems with LLMs, including LangChain, LangGraph, AWS Bedrock, Anthropic Claude, or OpenAI APIs. Candidates with classical ML backgrounds (PyTorch, scikit‑learn) who moved into generative AI architectures are welcome
- Hands‑on experience with LangGraph or similar stateful agent frameworks, particularly managing state graphs, loops, checkpointing, and human‑in‑the‑loop steps
- 3 to 4 years of React and TypeScript experience, enough to review UI code, understand frontend state, and ensure streaming APIs work cleanly with the web client
- Production AWS experience, including EKS or ECS, Lambda, Bedrock, S3, Docker, and Terraform or OpenTofu
- Authentication and security fundamentals, specifically OAuth 2.0, OIDC, and secure credential handling
- Nice to have Practical experience building custom MCP servers and defining tool discovery patterns
- Experience migrating from older retrieval systems (such as Amazon Kendra) to vector platforms like OpenSearch, Pinecone, or Qdrant
- Experience handling multimodal data (images, documents) within LangGraph state
- Background in Identity and Access Management (IAM) or Customer Identity (CIAM)