American Express is building foundational agentic AI platform capabilities for the enterprise, and this role helps design and deliver the core platform engineering that makes agents reliable, governed, observable, and easy to operate across multiple AI ecosystems. The work spans agent runtime and execution, a unified control plane, enterprise governance and security, discovery catalogs, end-to-end telemetry, evaluation and continuous learning, and AI CI/CD practices. This is an onsite opportunity in Sunrise, FL, with a salary range of USD 123,000 - 215,250 per year.
The platform is intended to help teams move from agent ideas to secure, repeatable production deployments through consistent self-service capabilities. The goal is to make the governed path the default path, so developers can build once and operate agents across heterogeneous execution environments and providers with the right controls, identity, and operational instrumentation included.
Responsibilities
- Contribute to the architecture and implementation of capabilities across the Agentic AI Platform, including Agent Runtime & Execution.
- Design and build scalable agent runtime infrastructure using Kubernetes and cloud-native technologies.
- Develop runtime abstractions that enable agents to execute in internally managed sandboxed environments and third-party AI platforms.
- Build capabilities for agent orchestration, tool execution, state and context management, memory, asynchronous workloads, event-driven execution, and multi-agent workflows.
- Build a unified control plane for managing agents across heterogeneous execution environments and AI providers.
- Develop APIs and services for agent registration, configuration, deployment, versioning, lifecycle management, policy enforcement, and runtime management.
- Engineer enterprise platform controls for AI governance, including agent identity, authentication and authorization, tool permissions, policy enforcement, data boundaries, auditability, and lifecycle controls.
- Build mechanisms for governing models, prompts, tools, MCP servers, knowledge sources, agent-to-agent interactions, and external integrations.
- Partner with security, risk, privacy, and governance teams to translate enterprise requirements into scalable technical controls.
- Create registries and catalogs for discovery of reusable agents, tools, skills, prompts, models, knowledge sources, and other platform capabilities.
- Develop metadata, ownership, versioning, dependency, certification, and discovery mechanisms for a healthy enterprise agent ecosystem.
- Build end-to-end telemetry for agent execution, including traces, events, model interactions, tool calls, latency, token consumption, cost, failures, policy decisions, and quality signals.
- Enable explainability and debuggability for complex and multi-agent workflows, helping platform and application teams understand behavior across models, runtimes, tools, and external systems.
- Develop evaluation frameworks and automated evaluation pipelines covering offline evaluations, production signals, human feedback, regression testing, and experimentation.
- Support continuous learning loops that turn production telemetry and feedback into improvements to agents, prompts, tools, models, and platform capabilities.
- Build SDKs, APIs, CLIs, templates, local development environments, and self-service workflows that make agent development simple and productive.
- Create opinionated paved roads that incorporate enterprise security, governance, observability, and operational standards.
- Build AI agent CI/CD capabilities including automated evaluation, policy validation, security checks, artifact/version management, deployment, promotion, rollback, and release controls, and apply GitOps and infrastructure-as-code patterns.
- Help establish engineering standards for moving agents from experimentation to production safely and repeatedly.
Requirements
- Strong software engineering experience building production systems using languages such as Python, Java, Go, or TypeScript.
- Strong understanding of Generative AI, LLMs, agent architectures, tool/function calling, retrieval, context management, and agent orchestration.
- Experience building production applications or platforms using major model providers or AI platforms.
- Experience with Kubernetes, containers, microservices, distributed systems, and cloud-native architecture.
- Experience designing production APIs and event-driven or asynchronous systems.
- Strong understanding of modern cloud infrastructure and infrastructure-as-code practices.
- Experience with CI/CD, automated testing, production observability, and software delivery practices.
- Strong understanding of security fundamentals including identity, authentication, authorization, secrets, and least-privilege access.
- Ability to navigate ambiguous technical problems and turn emerging technologies into reliable production systems.
- Strong communication skills and ability to collaborate across engineering, architecture, product, security, and governance organizations.
- Bachelor’s degree required.
Technologies
Python, Java, Go, TypeScript, Kubernetes, cloud-native technologies, CI/CD, GitOps, infrastructure-as-code, OpenTelemetry, OpenAI Agents SDK, LangGraph, Semantic Kernel, Google ADK, Model Context Protocol (MCP).
Preferred Qualifications
- Experience with agent frameworks and orchestration technologies such as LangGraph, Semantic Kernel, Google ADK, OpenAI Agents SDK, or similar frameworks.
- Model Context Protocol (MCP) and emerging agent interoperability protocols like A2A.
- Multi-agent systems and agent-to-agent communication patterns.
- Kubernetes operators, controllers, service meshes, or sophisticated Kubernetes platform engineering.
- Agent/model gateways and intelligent model routing.
- Vector databases, retrieval systems, embeddings, and enterprise knowledge architectures.
- LLM and agent evaluation frameworks, LLM-as-judge techniques, human feedback systems, experimentation, and quality measurement.
- AI observability, distributed tracing, OpenTelemetry, Langfuse and production monitoring.
- Experience with Evals.
- Policy engines and policy-as-code.
- AI security, prompt injection defenses, tool security, data-loss prevention, and AI-specific threat modeling.
- Platform engineering, internal developer platforms, developer portals, and enterprise service catalogs.
- Large-scale distributed systems operating in highly regulated environments.