AI Platform SRE Lead: Reliability & Scale for LLM Infra

Manychat

Barcelona

Híbrido

EUR 90.000 - 130.000

Jornada completa

Hace 7 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Relocation support
Health insurance for you and family
Professional development budget
Flexible benefits
Hybrid work model
In-office meals & snacks

Descripción de la vacante

Manychat, a leading chat marketing platform, is seeking an AI Platform SRE to own the reliability, performance, and cost of our AI infrastructure and to raise the bar for how our engineering org builds on LLMs.

You will shape the AI Gateway, design routing and failover, run capacity planning and incident response, and build strong observability with Prometheus, Grafana, and OpenTelemetry, partnering directly with the Head of Infrastructure.

Formación

  • 5+ years in SRE / platform / infra with production ownership.
  • Hands-on experience operating LLM-backed systems in production.
  • Cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD.
  • Strong observability practice and ability to define SLOs for non-deterministic systems.
  • Proven cost-optimization work and making spend visible.
  • Staff-level influence: set technical direction beyond your own team.

Responsabilidades

  • Own reliability and performance of our AI infrastructure: AI Gateway, inference services, and integrations with Bedrock, OpenAI, Anthropic, or similar.
  • Design and evolve the AI Gateway: routing, failover, rate limiting, caching, and guardrails.
  • Build observability for AI systems: latency/throughput/SLOs, token metrics, and drift signals.
  • Drive cost optimization and FinOps for AI workloads: visibility, right-sizing, caching strategies.
  • Run capacity planning and incident response for inference services; write/runbooks and postmortems.
  • Scale AI expertise across the org: set standards, review designs, and coach teams shipping LLM-backed features.

Conocimientos

SRE / infra
LLM production
Cloud-native
Observability
Cost optimization
Technical leadership

Herramientas

Prometheus
Grafana
OpenTelemetry

Descripción del empleo

Manychat, a leading chat marketing platform, is seeking an AI Platform SRE to own the reliability, performance, and cost of our AI infrastructure and to raise the bar for how our engineering org builds on LLMs.

You will shape the AI Gateway, design routing and failover, run capacity planning and incident response, and build strong observability with Prometheus, Grafana, and OpenTelemetry, partnering directly with the Head of Infrastructure.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

AI Platform Engineer: Scalable LLM Infra
AI Platform Engineer: Scalable LLM Infra

Peak3 (formerly ZA Tech) • Madrid

Presencial
EUR 70.000 - 100.000
Senior SRE: AI Platform & Cloud Reliability
Senior SRE: AI Platform & Cloud Reliability

Doist • Barcelona

Híbrido
EUR 85.000 - 120.000
Hybrid onboarding
Health insurance
Professional development budget
+4
Staff Site Reliability Engineer — AI Platform
Staff Site Reliability Engineer — AI Platform

Manychat • Barcelona

Híbrido
EUR 90.000 - 130.000
Relocation support
Health insurance for you and family
Professional development budget
+3
Remote Backend Lead — AI Platform & LLM Systems
Remote Backend Lead — AI Platform & LLM Systems

Acclaim AI • Barcelona

Presencial
EUR 90.000 - 130.000
Fully remote across Europe
Private English lessons via Preply
Company-paid subscriptions to top AI‑m
Remote AI Platform Engineer: Production-Ready LLM Infra
Remote AI Platform Engineer: Production-Ready LLM Infra

ProducePay • España

Presencial
EUR 85.000 - 130.000
Paid time off (40 days/year)
Wellbeing support
Wfh stipend
+1
AI Platform Engineer
AI Platform Engineer

Peak3 (formerly ZA Tech) • Madrid

Presencial
EUR 70.000 - 100.000
AI Platform Engineer—MLOps & Cloud Reliability
AI Platform Engineer—MLOps & Cloud Reliability

Verisure • Elche

Presencial
EUR 50.000 - 90.000
Remote-First Senior SRE: Scalable AI Infra
Remote-First Senior SRE: Scalable AI Infra

Runware • España

Presencial
EUR 70.000 - 100.000
Generous paid time off
Stock options
Remote-first setup
+3
AI Platform & Infra Engineer — Remote, Impactful ML Ops
AI Platform & Infra Engineer — Remote, Impactful ML Ops

Axiomatic_AI • Barcelona

Híbrido
EUR 70.000 - 90.000
Senior Platform Engineer — AI-First Infra & Scale
Senior Platform Engineer — AI-First Infra & Scale

Haddock • Barcelona

Híbrido
EUR 50.000 - 60.000
Hybrid work
Office in Barcelona Poblenou
AI tooling access
+2