We are looking for a Lead Platform Engineer to build and operate shared AI platform services that power agentic workflows, including LLM gateways, MCP tooling, and observability by default. You will enable teams with secure, scalable Kubernetes-based delivery, deep telemetry, and cost visibility.ResponsibilitiesBuild shared AI platform services such as LLM gateways, proxy layers, routing, fallback, and supporting APIs with telemetry and cost attributionOperate and scale platform services on Kubernetes using GitOps workflows with Helm and Kustomize, including progressive delivery and autoscalingOnboard model providers and model versions across environments with routing rules, quotas, tiering, fallback behavior, and deprecation pathsDesign and run the MCP layer, including server deployment, tool registration and discovery, session handling, and safe permission boundariesImplement authentication, authorization, and tenancy using OAuth2/OIDC, SSO, JWT, API keys, and managed secretsInstrument AI platform behavior end to end using OpenTelemetry, metrics, and structured logs alongside application-level AI tracesImplement Langfuse telemetry patterns for traces, spans, prompts, completions, feedback, evaluations, latency, errors, and token usageBuild dashboards, alerts, and reports for AI reliability, performance, quality, and evaluation outcomesDesign cost and usage observability across LLM vendors with attribution by app, team, user, model, workflow, and environmentCreate showback or chargeback-ready metrics and feed insights into routing, tiering, and capacity decisionsApply policy guardrails including PII detection, filtering, audit logging, and retention controls for prompts, completions, and tracesEnable the agent lifecycle with publishing, versioning, registry discovery, memory management, resilience, and explainability signalsBuild self-service onboarding and provisioning workflows, templates, and portals or CLIs for teamsPartner with engineering teams to define and enforce observability standards as the default delivery pathRequirements5+ years platform engineering experienceStrong leadership skills to set standards and mentor engineersProven project ownership for designing and operating shared internal platformsAdvanced Kubernetes skills for running production servicesStrong AI platforms experience across model gateways, routing, and agent runtimesHands-on Langfuse experience for AI tracing and evaluation metadataPractical Model Context Protocol (MCP) experience with servers, registries, and tool boundariesStrong OpenTelemetry skills for traces, metrics, and structured logsSolid security knowledge in OAuth2/OIDC, SSO, JWT, and secrets managementExcellent stakeholder communication skills with application and developer teamsUpper-Intermediate English proficiency (B2)Nice to haveLangChain experience for agent workflow integration and analysisLangGraph experience for graph-based agent orchestrationRetrieval-Augmented Generation (RAG) experience including evaluation and tuningWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.