Senior SRE Engineer

Codeway

Barcelona

Presencial

EUR 85.000 - 120.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Codeway is seeking a Senior Site Reliability Engineer to own reliability, performance, and security across our global platform. You will define SLIs/SLOs, own observability, manage multi‑cluster Kubernetes, and drive incident response with a security‑first mindset.

You’ll partner with product engineering to build self‑service tooling, optimize capacity and cloud cost, and embed security into architecture and operations as we scale. This is a hands‑on role demanding ownership and collaboration.

Formación

  • Experience operating high-traffic production systems.
  • Hands-on Kubernetes experience in production.
  • Track record with SLOs and error budgets.
  • Ability to communicate with engineers and leadership under pressure.

Responsabilidades

  • Define, instrument, and report on SLIs, SLOs, and error budgets across critical services.
  • Own observability end‑to‑end — metrics, logs, traces, dashboards, and alerting.
  • Reduce alert noise and false positives so on‑call engineers can trust what wakes them up.
  • Run reliability reviews and an error‑budget policy that shapes how teams prioritize between shipping and stability.
  • Operate, scale, and upgrade multi‑cluster Kubernetes (GKE) environment.
  • Build self‑service platform tooling that lets product teams move quickly.
  • Embed security into the platform through RBAC, secrets management, and patch cadence.
  • Lead on-call rotation and incident commander during production incidents.
  • Develop and maintain Terraform IaC and CI/CD pipelines; enforce GitOps and progressive delivery.
  • Eliminate operational toil through automation and tooling.

Conocimientos

SRE experience
Kubernetes
Cloud engineering
SLOs & budgets
IaC
CI/CD
Observability
Security mindset
Automation scripting
Incident command
Communication skills

Herramientas

Terraform
GitOps
CI/CD tooling
GKE

Descripción del empleo

ABOUT CODEWAY

Codeway is a global consumer tech company with more than 400M users worldwide.

Since 2020, we’ve built and scaled 60+ mobile apps across creativity, productivity, wellness, language learning, and entertainment.

Our flagship apps — Retake AI, Cleanup, Learna, and DramaPops — and many of them lead their categories globally. In 2024, we became the most downloaded app publisher on iOS, driven by cutting‑edge AI research, sharp data‑driven execution, and a relentless focus on product and marketing.

We’re a team of 300+ people across İstanbul and Barcelona who bring curiosity, passion, trust, and ownership to everything we build. Recognized as a #1 LinkedIn Top Startup and a Great Place to Work in Europe, Codeway is where ambitious people do their life’s best work.

We’re building the next generation of consumer tech and reimagining what mobile apps can be.

This is Codeway. This is our way. Join us.

POSITION

We’re looking for a Senior Site Reliability Engineer to own and mature reliability, performance, and security across our growing platform. This role sits at the intersection of Engineering, Infrastructure, and Security, helping design, operate, and continuously improve the systems that keep dozens of consumer apps running for users around the world.

You’ll work closely with product engineering teams to make reliability measurable rather than assumed. That means defining and enforcing SLIs, SLOs, and error budgets; operating and hardening a multi‑cluster Kubernetes environment; building the observability that catches problems before users feel them; and leading incident response when things break. It’s a hands‑on role with real ownership over how reliability and security evolve as we scale.

Several parts of our reliability practice are still early in their maturity. We’re looking for someone who enjoys building the standards, processes, tooling, and automation that will form the foundation of our SRE function — not someone waiting for a playbook to already exist.

We welcome applicants from all backgrounds and experiences. If you’re excited about running systems at consumer scale and believe you could be a strong fit, we encourage you to apply, even if your experience doesn’t align perfectly with every qualification listed below.

WHAT YOU’LL BE DOING Reliability, SLOs, Observability
  • Define, instrument, and report on SLIs, SLOs, and error budgets across critical services, so reliability decisions are driven by data rather than opinion.

  • Own observability end‑to‑end — metrics, logs, traces, dashboards, and alerting — and drive measurable reductions in detection and resolution times.

  • Reduce alert noise and false positives so on‑call engineers can trust what wakes them up.

  • Run reliability reviews and an error‑budget policy that shapes how teams prioritize between shipping and stability.

Kubernetes & Platform Operations
  • Operate, scale, and upgrade our multi‑cluster Kubernetes (GKE) environment: cluster lifecycle, autoscaling, networking, ingress, and resource management.

  • Act as the deep‑expertise escalation point for cluster and platform issues across dozens of services.

  • Own capacity planning, performance, and cloud cost efficiency, balancing spend against reliability targets.

  • Build self‑service platform tooling that lets product teams move quickly without needing to become infrastructure experts.

Security & Resilience
  • Embed security into the platform through RBAC and least‑privilege, secrets management, image and dependency scanning, network policies, and a disciplined patching cadence.

  • Partner with the security function on vulnerability remediation, audit readiness, and secure‑by‑default infrastructure.

  • Own disaster recovery: define and regularly validate RTO/RPO targets through DR drills and failure testing.

  • Contribute to architecture and production‑readiness reviews so reliability and security are designed in, not bolted on.

Incident Response & Automation
  • Lead the on‑call rotation and act as incident commander during production incidents.

  • Run blameless postmortems, quantify impact, and track corrective actions through to closure so the same failure doesn’t recur.

  • Build and maintain Infrastructure as Code (Terraform) and CI/CD pipelines, enforcing GitOps and progressive delivery with automated rollbacks.

  • Systematically identify, measure, and eliminate operational toil through automation, protecting engineering time for high‑leverage work.

WHAT YOU’LL BRING?
  • Experience operating high‑traffic, always‑on production systems at meaningful scale, typically gained over 5–8 years in SRE, Platform, or DevOps roles.

  • Hands‑on production Kubernetes experience — you’ve run clusters day to day, through upgrades, autoscaling, and real troubleshooting under load, not just deployed to them.

  • A strong cloud engineering background, along with solid Linux and networking fundamentals.

  • A track record of defining and operating with SLOs and error budgets, and comfort being measured on reliability outcomes.

  • Experience with Infrastructure as Code and CI/CD pipeline design — you treat infrastructure and delivery as code.

  • Depth in observability tooling: instrumentation, dashboarding, and alert design.

  • A genuine security‑first mindset, where least‑privilege, secrets hygiene, and vulnerability management are habits rather than afterthoughts.

  • Scripting and automation fluency in at least one language, used to build tooling and remove toil.

  • Incident‑command experience: owning on‑call, running blameless postmortems, and driving resolution times down over time.

  • Ability to communicate clearly with both engineers and leadership, especially under pressure.

NICE TO HAVE
  • Experience with high‑scale consumer or mobile app backends, or with AI/ML inference workloads and their scaling characteristics.

  • Experience with GitOps and progressive‑delivery patterns such as canary and blue‑green rollouts.

  • Familiarity with service mesh, API gateways, or multi‑region and multi‑cluster topologies.

  • Cloud cost optimization and FinOps discipline at scale.

  • Exposure to compliance initiatives (SOC 2, ISO 27001, GDPR) and broader DevSecOps practice.

  • Chaos engineering or resilience testing experience.

  • Relevant certifications in Kubernetes, cloud, or DevOps disciplines.

  • Experience supporting many independent services and teams concurrently in a fast‑shipping, product‑led environment.

OUR ENVIRONMENT

You’ll help operate and improve a modern, cloud‑native environment built around:

  • Kubernetes (GKE) and containerized workloads

  • Google Cloud Platform (GCP)

  • Terraform and Infrastructure as Code

  • CI/CD and GitOps tooling

  • Modern observability (metrics, logs, traces, alerting)

  • Cloudflare CDN and edge

Experience with these exact platforms is beneficial but not required. We value strong fundamentals, curiosity, and the ability to quickly learn new technologies and environments.

WHAT SUCCESS LOOKS LIKE

Within your first 12 months, you’ll help establish and mature key reliability capabilities, including:

  • Clear SLIs, SLOs, and error budgets live and reported for our most critical services.

  • A measurable reduction in detection and resolution times for production incidents.

  • A consistent, blameless incident response practice, with postmortem actions tracked to closure.

  • A hardened Kubernetes fleet, with key security gaps closed across access control, secrets, scanning, and patching.

  • Expanded Infrastructure as Code and observability coverage across the platform.

  • Disaster recovery drills that reliably pass agreed RTO/RPO targets.

  • Operational toil trending down against an explicit target, with automation replacing manual work.

  • Product teams self

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Randstad (Schweiz) AG • Madrid

Presencial
EUR 70.000 - 100.000
Health Insurance
Paid Time Off (PTO)
Paid Holidays
+2
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Ribarroja del Turia

Híbrido
EUR 40.000 - 70.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps
Competitive compensation: USD-based pay with education, fitness, and team activity budgets
Exciting projects: Modern solutions with Fortune 500 and top product companies
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Madrid

Híbrido
EUR 45.000 - 60.000
Professional growth
Competitive compensation
Exciting projects
+1
Operations Manager
Operations Manager

Destiny group • Ribadeo

Híbrido
EUR 90.000 - 120.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

F. Hoffmann-La Roche AG • Sant Cugat del Vallès

Presencial
EUR 50.000 - 75.000
Site Reliability Engineer - DevSecOps Engineer
Site Reliability Engineer - DevSecOps Engineer

Kyndryl • España

Híbrido
EUR 65.000 - 90.000
Senior Site Reliability Engineer — Reliability & Platform Ops
Senior Site Reliability Engineer — Reliability & Platform Ops

Codeway • Barcelona

Presencial
EUR 85.000 - 120.000
Senior DevOps / Site Reliability Engineer
Senior DevOps / Site Reliability Engineer

N-iX • España

Híbrido
EUR 60.000 - 100.000
Flexible work format
Competitive salary
Education reimbursement
+2
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Camlin Group • Málaga

Presencial
EUR 85.000 - 120.000
Staff Software Engineer, Service Infrastructure
Staff Software Engineer, Service Infrastructure

United States Digital Space LLC • Amer

Híbrido
EUR 130.000 - 190.000