Lead Platform Engineering

EPAM Systems, Inc.

Brasil

Teletrabalho

BRL 400 000 - 520 000

Tempo integral

Há 5 dias
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Uma candidatura completa num minuto — currículo personalizado e carta de apresentação, prontos a enviar.

Ultrapassa os filtros ATS

Vantagens oferecidas por esta oferta de emprego

Healthcare benefits
Paid time off and sick leave
Upskilling, reskilling and certificate
LinkedIn Learning access
Global career opportunities
Volunteer and community involvement

Resumo da oferta

EPAM Systems, Inc. is seeking a Lead Platform Engineer to design and operate shared AI platform services on Kubernetes. You will implement authentication, telemetry, and cost attribution while enabling teams with scalable, secure delivery.

This role emphasizes Langfuse-based AI tracing, observability, and GitOps with Helm and Kustomize across environments. You will onboard model providers, manage quotas and routing, and partner with teams to enforce standards and guardrails for security and data

Qualificações

  • 5+ years platform engineering experience with internal shared platforms.
  • Strong Kubernetes production experience.
  • Hands-on Langfuse experience for AI tracing and evaluation metadata.
  • Security and secrets management expertise (OAuth2/OIDC, SSO, JWT).
  • Excellent stakeholder communication with application and developer teams.
  • Upper-Intermediate English proficiency (B2).

Responsabilidades

  • Build shared AI platform services such as LLM gateways, proxy layers, routing, fallback, and supporting APIs with telemetry and cost attribution.
  • Operate and scale platform services on Kubernetes using GitOps workflows with Helm and Kustomize, including progressive delivery and autoscaling.
  • Onboard model providers and model versions across environments with routing rules, quotas, tiering, fallback behavior, and deprecation paths.
  • Design and run the MCP layer, including server deployment, tool registration and discovery, session handling, and safe permission boundaries.
  • Implement authentication, authorization, and tenancy using OAuth2/OIDC, SSO, JWT, API keys, and managed secrets.
  • Instrument AI platform behavior end to end using OpenTelemetry, metrics, and structured logs alongside application-level AI traces.
  • Langfuse telemetry patterns for traces, spans, prompts, completions, feedback, evaluations, latency, errors, and token usage.
  • Build dashboards, alerts, and reports for AI reliability, performance, quality, and evaluation outcomes.
  • Design cost and usage observability across LLM vendors with attribution by app, team, user, model, workflow, and environment.
  • Create showback or chargeback-ready metrics for routing, tiering, and capacity decisions.
  • Apply policy guardrails including PII detection, filtering, audit logging, and retention controls for prompts, completions, and traces.
  • Enable the agent lifecycle with publishing, versioning, registry discovery, memory management, resilience, and explainability signals.
  • Build self-service onboarding and provisioning workflows, templates, and portals or CLIs for teams.
  • Partner with engineering teams to define and enforce observability standards as the default delivery path.

Conhecimentos

Kubernetes
GitOps
Helm
Kustomize
Langfuse
OAuth2/OIDC
SSO
JWT
Observability
OpenTelemetry
LangChain
RAG
English (B2)

Ferramentas

Helm
Kustomize
Langfuse
LangChain

Descrição da oferta de emprego

We are looking for a Lead Platform Engineer to build and operate shared AI platform services that power agentic workflows, including LLM gateways, MCP tooling, and observability by default. You will enable teams with secure, scalable Kubernetes-based delivery, deep telemetry, and cost visibility.ResponsibilitiesBuild shared AI platform services such as LLM gateways, proxy layers, routing, fallback, and supporting APIs with telemetry and cost attributionOperate and scale platform services on Kubernetes using GitOps workflows with Helm and Kustomize, including progressive delivery and autoscalingOnboard model providers and model versions across environments with routing rules, quotas, tiering, fallback behavior, and deprecation pathsDesign and run the MCP layer, including server deployment, tool registration and discovery, session handling, and safe permission boundariesImplement authentication, authorization, and tenancy using OAuth2/OIDC, SSO, JWT, API keys, and managed secretsInstrument AI platform behavior end to end using OpenTelemetry, metrics, and structured logs alongside application-level AI tracesImplement Langfuse telemetry patterns for traces, spans, prompts, completions, feedback, evaluations, latency, errors, and token usageBuild dashboards, alerts, and reports for AI reliability, performance, quality, and evaluation outcomesDesign cost and usage observability across LLM vendors with attribution by app, team, user, model, workflow, and environmentCreate showback or chargeback-ready metrics and feed insights into routing, tiering, and capacity decisionsApply policy guardrails including PII detection, filtering, audit logging, and retention controls for prompts, completions, and tracesEnable the agent lifecycle with publishing, versioning, registry discovery, memory management, resilience, and explainability signalsBuild self-service onboarding and provisioning workflows, templates, and portals or CLIs for teamsPartner with engineering teams to define and enforce observability standards as the default delivery pathRequirements5+ years platform engineering experienceStrong leadership skills to set standards and mentor engineersProven project ownership for designing and operating shared internal platformsAdvanced Kubernetes skills for running production servicesStrong AI platforms experience across model gateways, routing, and agent runtimesHands-on Langfuse experience for AI tracing and evaluation metadataPractical Model Context Protocol (MCP) experience with servers, registries, and tool boundariesStrong OpenTelemetry skills for traces, metrics, and structured logsSolid security knowledge in OAuth2/OIDC, SSO, JWT, and secrets managementExcellent stakeholder communication skills with application and developer teamsUpper-Intermediate English proficiency (B2)Nice to haveLangChain experience for agent workflow integration and analysisLangGraph experience for graph-based agent orchestrationRetrieval-Augmented Generation (RAG) experience including evaluation and tuningWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

Senior Platform Engineer (AI Platforms)
Senior Platform Engineer (AI Platforms)

EPAM Systems, Inc. • Brasil

Teletrabalho
BRL 180 000 - 300 000
Healthcare benefits
Paid time off
Upskilling programs
+2
Lead AI Engineer
Lead AI Engineer

EPAM Systems, Inc. • Brasil

Teletrabalho
BRL 300 000 - 420 000
Lead AI Agentic Developer
Lead AI Agentic Developer

EPAM Systems, Inc. • Brasil

Teletrabalho
BRL 350 000 - 520 000
Lead Python AI Engineer
Lead Python AI Engineer

EPAM Systems, Inc. • Brasil

Teletrabalho
BRL 180 000 - 340 000
Healthcare benefits
Paid time off
Upskilling & certification
+2
Senior AI Engineer (Python)
Senior AI Engineer (Python)

EPAM Systems, Inc. • Brasil

Híbrido
BRL 180 000 - 240 000
Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
+1
Senior Fullstack Developer
Senior Fullstack Developer

EPAM Systems, Inc. • Brasil

Teletrabalho
BRL 180 000 - 300 000
Healthcare benefits
Paid time off
Upskilling and certification courses
+2
Senior AI Engineer
Senior AI Engineer

EPAM Systems, Inc. • Brasil

Teletrabalho
BRL 180 000 - 320 000
Lead AWS DevOps Engineer
Lead AWS DevOps Engineer

EPAM Systems, Inc. • Brasil

Teletrabalho
BRL 180 000 - 320 000
Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
+1
Lead Forward Deployed Engineer
Lead Forward Deployed Engineer

EPAM Systems, Inc. • Brasil

Teletrabalho
BRL 180 000 - 300 000
AI Center of Excellence (COE) Engineering Manager
AI Center of Excellence (COE) Engineering Manager

EPAM Systems, Inc. • Brasil

Teletrabalho
BRL 300 000 - 520 000