Staff Site Reliability Engineer — AI Platform

Manychat

Bellprat

Híbrido

EUR 90.000 - 130.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Hybrid onboarding
Health insurance for you and family
Development budget
Flexible benefits package
Hybrid work schedule
In-office meals and snacks

Descripción de la vacante

Manychat is seeking an AI-native Site Reliability Engineer to own the reliability, performance, and cost of our AI infrastructure from day one. You will design the AI Gateway, operate AI services across Bedrock and OpenAI providers, and build observability, SLOs, and runbooks that scale with a fast-growing product used by customers worldwide.

Hybrid work, relocation support, and a competitive compensation package are part of the offer, with direct collaboration to the Head of Infrastructure and

Formación

  • 5+ years in SRE / platform / infrastructure engineering with production ownership at scale.
  • Hands-on experience operating LLM-backed systems in production.
  • Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD.
  • Strong observability practice and experience defining SLOs for non-deterministic systems.
  • Proven cost-optimization work with visible impact.
  • Staff-level influence and ability to shape technical direction beyond own team.

Responsabilidades

  • Own reliability and performance of AI infrastructure including AI Gateway and inference services.
  • Design and evolve the AI Gateway with routing, failover, rate limiting, caching and guardrails.
  • Build observability for AI systems with latency, throughput, SLOs, and token-level metrics.
  • Drive cost optimization and FinOps for AI workloads with visibility and right-sizing.
  • Run capacity planning and incident response for inference services; write postmortems.
  • Scale AI expertise across the org by setting standards and coaching teams.

Conocimientos

SRE experience
LLM in production
Cloud-native (AWS)
Kubernetes
Terraform/IaC
Observability (SLOs)
Cost optimization
Technical leadership

Herramientas

Prometheus
Grafana
OpenTelemetry
Kubernetes
Terraform

Descripción del empleo

WHO WE ARE

Manychat is a leading chat marketing platform, helping businesses engage with their customers on Instagram, Facebook Messenger, WhatsApp, and Telegram. Trusted by over 1 million brands in 170+ countries, we're an official Meta Business Partner, backed by Bessemer Venture Partners. With 400+ teammates across offices in Austin, Barcelona, Yerevan, São Paulo, and Amsterdam — we're building the infrastructure of conversational commerce.

WHO WE’RE LOOKING FOR

Want to shape an AI platform's architecture from day one, instead of just maintaining what someone else built?

Manychat runs AI features for businesses worldwide, and we're hiring an SRE to own the reliability, performance, and cost of that AI infrastructure, and to raise the bar for how our whole engineering org builds on LLMs.

This isn't a classical SRE role with a bit of AI sprinkled on top. We need someone AI-native: you already understand how modern LLM systems behave in production, things like token throughput, provider rate limits, degraded model quality, and inference latency tails, and you treat all of that as a first-class reliability concern, not an afterthought.

Why this role is worth your time:

  • The AI Platform is still young, so you'll shape its architecture, standards, and roadmap from the ground up.
  • The stakes are real. AI features sit right in the critical path of customer-facing automation, so your work actually matters to the business, not just to a dashboard.
  • You'll partner directly with the Head of Infrastructure, with real autonomy and visibility into how your decisions play out.

Sound like the kind of ownership you've been looking for?

WHAT YOU'LL DO
  • Own reliability and performance of our AI infrastructure: AI Gateway, inference services, and integrations with Amazon Bedrock, Azure OpenAI, and other LLM providers.
  • Design and evolve the AI Gateway: routing, failover between providers, rate limiting, caching, and guardrails.
  • Build observability for AI systems: latency/throughput/error SLOs per model and provider, token-level metrics, quality and drift signals.
  • Drive cost optimization and FinOps for AI workloads: per-feature cost visibility, model right-sizing, caching strategies, provider mix.
  • Run capacity planning and incident response for inference services; write and improve runbooks and postmortems.
  • Scale AI expertise across the org: set standards, review designs, and coach teams shipping LLM-backed features.
TO SHINE IN THIS ROLE

You’ll need:

  • 5+ years in SRE / platform / infrastructure engineering, including production ownership at significant scale.
  • Hands‑on experience operating LLM‑backed systems in production: provider APIs (Bedrock, OpenAI, Anthropic, or similar), inference pipelines, self‑hosted or managed model serving.
  • Deep cloud‑native background: AWS, Kubernetes, Terraform/IaC, CI/CD.
  • Strong observability practice (Prometheus/Grafana, OpenTelemetry, or equivalent) and experience defining SLOs for non‑deterministic systems.
  • Proven cost‑optimization work: you can show where you cut cloud or inference spend and how you made cost visible.
  • Staff‑level influence: you've set technical direction beyond your own team and brought others along

It would be great if you have:

  • Experience building or operating an LLM gateway/proxy (e.g., LiteLLM, Kong AI Gateway, custom).
  • Experience with GPU workload optimization, quantization, or serving frameworks (vLLM, TGI, Triton).
  • Experience with eval pipelines and quality monitoring for LLM outputs.
Why this role
  • Green‑field ownership: the AI Platform is young; you'll shape its architecture, standards, and roadmap.
  • Real scale and real stakes: AI features sit in the critical path of customer‑facing automation.
  • Direct partnership with the Head of Infrastructure; high autonomy and visibility.
WHAT WE OFFER

We care deeply about your growth, well‑being, and comfort:

  • Hybrid onboarding to start work remotely and relocation support for you and your family.
  • Comprehensive health insurance for both you and your family.
  • Professional development budget for conference tickets, online courses, and other relevant resources to help you grow.
  • Flexible benefits package: no one‑size‑fits‑all perks here. You get a budget and you decide where it goes, from health and wellbeing to family and setting up your home office.
  • Hybrid work and generous, flexible time off — planned with your team, not rationed by a rigid quota.
  • In‑office perks, including free meals and snacks.
  • Company‑funded sport activities, annual offsites and team‑building events.
  • AI isn't a perk here — it's how we work. Claude, OpenAI, and more are on by default from day one, and we back teams in adopting whatever makes them faster.

Manychat is an Equal Opportunity Employer. We’re committed to building a diverse and inclusive team. We do not discriminate against qualified employees or applicants because of race, color, religion, gender identity, sex, sexual preference, sexual identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical condition, military status, or any other characteristic protected by local law or ordinance.

This commitment is also reflected through our candidate experience. If you have individual needs that may require an accommodation during the interview process, please indicate this in your application. We will do our best to provide assistance throughout your interview process to ensure you’re set up for success.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Staff Site Reliability Engineer (AI Platform)
Staff Site Reliability Engineer (AI Platform)

Manychat • Barcelona

Híbrido
EUR 85.000 - 120.000
Hybrid onboarding
Relocation support
Health insurance
+5
Staff Site Reliability Engineer — AI Platform
Staff Site Reliability Engineer — AI Platform

Manychat • Barcelona

Híbrido
EUR 90.000 - 130.000
Relocation support
Health insurance for you and family
Professional development budget
+3
Staff AI Platform SRE — Build Reliable LLM Infra
Staff AI Platform SRE — Build Reliable LLM Infra

Manychat • Barcelona

Híbrido
EUR 85.000 - 120.000
Hybrid onboarding
Relocation support
Health insurance
+5
Head of Workplace & Experience Barcelona, Spain
Head of Workplace & Experience Barcelona, Spain

ManyChat, Inc. • Barcelona

Híbrido
EUR 80.000 - 110.000
Hybrid onboarding
Comprehensive health insurance
Professional development budget
+3
Senior Python Engineer, Brands Barcelona, Barcelona, Spain
Senior Python Engineer, Brands Barcelona, Barcelona, Spain

ManyChat, Inc. • Barcelona

Híbrido
EUR 70.000 - 100.000
Hybrid onboarding
Relocation support
Health insurance for you and family
+7
Data Platform Engineer
Data Platform Engineer

Manychat • Bellprat

Híbrido
EUR 70.000 - 100.000
Hybrid onboarding
Relocation support
Health insurance
+5
AI Platform SRE Lead: Reliability & Scale for LLM Infra
AI Platform SRE Lead: Reliability & Scale for LLM Infra

Manychat • Barcelona

Híbrido
EUR 90.000 - 130.000
Relocation support
Health insurance for you and family
Professional development budget
+3
Senior AI Platform SRE: Reliability, Cost & Scale
Senior AI Platform SRE: Reliability, Cost & Scale

Manychat • Bellprat

Híbrido
EUR 90.000 - 130.000
Hybrid onboarding
Health insurance for you and family
Development budget
+3
Senior Python Engineer, Billing & Accounts
Senior Python Engineer, Billing & Accounts

Manychat • Bellprat

Híbrido
EUR 70.000 - 100.000
Hybrid onboarding, relocation support
Comprehensive health insurance
Professional development budget
+4
Senior Python Engineer (Billing)
Senior Python Engineer (Billing)

Manychat • Barcelona

Híbrido
EUR 70.000 - 110.000
Hybrid onboarding with remote start
Health insurance for you and family
Professional development budget
+4