Staff Site Reliability Engineer (AI Platform)

Manychat, Inc

Barcelona

Presencial

EUR 90.000 - 120.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Hybrid work model
Relocation support
Health insurance
Professional development budget
Team offsites

Descripción de la vacante

Manychat, Inc is seeking a Senior Site Reliability Engineer to join our hybrid team in Barcelona. You will shape reliability for our cloud-native platform, balancing hands-on ops with strategic improvements across Kubernetes, Terraform, and CI/CD pipelines.

You will collaborate with engineers to drive platform reliability and scale Kubernetes effectively. The role emphasizes owning infrastructure, securing cloud resources, and maintaining robust monitoring with Prometheus and Grafana, while

Formación

  • 5+ years of production Linux experience (Ubuntu/Amazon Linux).
  • Strong Kubernetes experience, ideally EKS, with Helm.
  • Comfort running Python workloads in containers and debugging.
  • Solid networking, IAM, and cloud security knowledge.
  • Hands-on with Nginx in ingress/reverse proxy scenarios.
  • Excellent written and verbal communication for devs.

Responsabilidades

  • Maintain and harden AWS infrastructure components (EC2, ALB/NLB, WAF, IAM, CloudWatch).
  • Operate and evolve EKS clusters powering Python-based AI services.
  • Migrate services to Kubernetes using Terraform and Helm.
  • Codify infrastructure with Terraform; manage host automation via Ansible.
  • Build and improve CI/CD pipelines with GitHub Actions.
  • Own observability: Prometheus, Grafana, alerts, on-call readiness.
  • Support OS patching, certs, WAF rules, and infra hygiene.
  • Create clear infra documentation and playbooks.
  • Occasionally support off-hours incidents (rare).

Conocimientos

5+ years SRE experience
Kubernetes (EKS)
Terraform
Helm
Docker / containers
Python workloads
Linux production
Networking & cloud security
Nginx (Ingress/reverse proxy)
CI/CD pipelines
Observability (Prometheus, Grafana)
Terraform and Ansible
Dev-friendly platform
Strong communication

Herramientas

Terraform
Helm
Ansible
GitHub Actions
Prometheus
Grafana
Nginx
AWS
EKS

Descripción del empleo

WHO WE ARE

We help creators get more out of every conversation with Instagram-focused automations and support for other channels like Messenger, WhatsApp, and TikTok. The result? Better engagement, more sales, and real, sustainable growth.

With a diverse team of 350+ people spread across three continents, we're building the leading Chat Marketing platform that is used — and loved — by more than 1.5 million customers worldwide.

WHO WE'RE LOOKING FOR

We're looking for a Senior Site Reliability Engineer who thrives at the crossroads of classic Linux and AWS infrastructure and modern Site Reliability Engineering. This is a high-impact, hybrid role designed for someone who can manage cloud resources, harden Kubernetes clusters, and shape a more reliable and developer-friendly platform.

We need you not just to maintain but to rethink and evolve our infrastructure, balancing hands-on operations with strategic improvements that future-proof our growing AI product landscape.

You'll take over key responsibilities from our current Infra Lead who is transitioning to a software-focused role, giving you immediate ownership and space to shine.

WHY THE ROLE IS SPECIAL

You won't be a cog in a massive SRE org. You'll be the bridge between Infrastructure and Engineering, shaping how we scale Kubernetes, how we approach platform reliability, and how developers ship fast without fear. You'll get autonomy, ownership, and a smart, humble team excited to learn with you.

WHAT YOU'LL DO

  • Maintain and harden AWS infrastructure (EC2, ALB/NLB, WAF, IAM, CloudWatch)
  • Operate and evolve our EKS clusters powering Python-based AI services
  • Migrate existing services to Kubernetes using Terraform and Helm
  • Codify infrastructure with Terraform and manage host-level automation via Ansible
  • Build and improve CI/CD pipelines with GitHub Actions
  • Own observability efforts: Prometheus, Grafana, alerting, and on-call readiness
  • Support OS-level patching, certs, WAF rules, and general infra hygiene
  • Partner with engineers to guide best practices and drive platform reliabilityCreate clean, maintainable infrastructure documentation and playbooks
  • Occasionally support rare off-hours incidents (don't worry, really rare)
TO SHINE IN THIS ROLE
  • 5+ years of experience managing Linux in production (Ubuntu, Amazon Linux)
  • Strong experience with Kubernetes (ideally EKS), Helm, and Terraform
  • Comfort with running and debugging Python workloads in containers
  • Solid understanding of networking, IAM, and cloud security best practices
  • Hands-on Nginx experience (Ingress and reverse proxy setups)
  • Excellent communication skills; you can explain complex infra to devs clearly
NICE TO HAVE SKILLS
  • Strong Ansible skills beyond the basics
  • PostgreSQL or Amazon RDS tuning and operations experience
  • Deep understanding of observability tools (Prometheus, Grafana, Loki, etc.)
  • Familiarity with PHP production environments
  • Experience with TDD, CI/CD best practices, and agile development
  • Any previous SRE-like exposure such as building resilience, automation, or incident tooling

WHAT WE OFFER
We care deeply about your growth, well-being, and comfort:

  • Hybrid onboarding to start work remote and relocation support for you and your family.
  • Comprehensive health insurance for both you and your family.
  • Professional development budget for conference tickets, online courses, and other relevant resources to help you grow.
  • Flexible benefits package to tailor perks that matters most for you.
  • Hybrid work and generous leave options to prioritize your work-life balance.
  • In-office perks, including free meals and snacks.
  • Company-funded sport activities, annual offsites and team-building events.

Manychat is an Equal Opportunity Employer. We're committed to building a diverse and inclusive team. We do not discriminate against qualified employees or applicants because of race, color, religion, gender identity, sex, sexual preference, sexual identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical condition, military status, or any other characteristic protected by local law or ordinance.

_This commitment is also reflected through our candidate experience. If you have individual needs that may require an accommodation during the interview process, please indicate this in your application. We will do our best to provide assistance throughout your interview process to ensure you're set up for success._

With my application, I accept the ManychatPrivacy Policy._

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Staff Site Reliability Engineer — AI Platform
Staff Site Reliability Engineer — AI Platform

Manychat • Bellprat

Híbrido
EUR 90.000 - 130.000
Hybrid onboarding
Health insurance for you and family
Development budget
+3
Staff Site Reliability Engineer (AI Platform)
Staff Site Reliability Engineer (AI Platform)

Manychat • Barcelona

Híbrido
EUR 85.000 - 120.000
Hybrid onboarding
Relocation support
Health insurance
+5
Data Platform Engineer
Data Platform Engineer

Manychat, Inc • Barcelona

Híbrido
EUR 60.000 - 90.000
Hybrid onboarding
Relocation support
Health insurance for you and family
+5
Data Platform Engineer
Data Platform Engineer

Manychat • Bellprat

Híbrido
EUR 70.000 - 100.000
Hybrid onboarding
Relocation support
Health insurance
+5
Senior Data Engineer (Data Platform)
Senior Data Engineer (Data Platform)

Manychat • Barcelona

Híbrido
EUR 40.000 - 70.000
Comprehensive health insurance
Professional development budget
Flexible benefits package
+2
Staff Site Reliability Engineer — AI Platform
Staff Site Reliability Engineer — AI Platform

Manychat • Barcelona

Híbrido
EUR 90.000 - 130.000
Relocation support
Health insurance for you and family
Professional development budget
+3
Senior Python Engineer (Billing)
Senior Python Engineer (Billing)

Manychat • Barcelona

Híbrido
EUR 70.000 - 110.000
Hybrid onboarding with remote start
Health insurance for you and family
Professional development budget
+4
Senior Python Engineer, Brands Barcelona, Barcelona, Spain
Senior Python Engineer, Brands Barcelona, Barcelona, Spain

ManyChat, Inc. • Barcelona

Híbrido
EUR 70.000 - 100.000
Hybrid onboarding
Relocation support
Health insurance for you and family
+7
Data Platform Engineer
Data Platform Engineer

Manychat • Barcelona

Híbrido
EUR 90.000 - 120.000
Relocation support
Health insurance
Dev budget for training
+5
Senior Python Engineer, Billing & Accounts
Senior Python Engineer, Billing & Accounts

Manychat • Bellprat

Híbrido
EUR 70.000 - 100.000
Hybrid onboarding, relocation support
Comprehensive health insurance
Professional development budget
+4