Senior SRE Engineer

Codeway

Barcelona

On-site

EUR 90,000 - 120,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Codeway in Barcelona is seeking a Senior Site Reliability Engineer to own reliability, performance, and security across our consumer platform. You’ll define SLIs, SLOs, error budgets, and operate a multi-cluster Kubernetes environment with observability at the core of decision-making.

You’ll partner with product engineering, security, and platform teams to automate, scale, and harden systems for global user growth.

Qualifications

  • Experience operating high-traffic, always-on production systems at scale.
  • Hands-on production Kubernetes experience and troubleshooting under load.
  • Strong cloud engineering background with Linux, networking basics.
  • Track record of defining and operating with SLOs and error budgets.
  • Infrastructure as Code and CI/CD pipeline design experience.

Responsibilities

  • Define, instrument, and report on SLIs, SLOs, and error budgets.
  • Own observability end-to-end and drive reduction in detection/resolution times.
  • Operate multi-cluster Kubernetes (GKE) environment and capacity planning.
  • Build self-service tooling and automation to reduce toil.
  • Lead on-call, incident response, blameless postmortems, and remediation.

Skills

SRE
Kubernetes
Cloud engineering
Linux networking
Observability
Scripting
Incident command
Communication

Tools

Terraform
CI/CD
GitOps

Job description

ABOUT CODEWAY

Codeway is a global consumer tech company with more than 400M users worldwide.

Since 2020, we’ve built and scaled 60+ mobile apps across creativity, productivity, wellness, language learning, and entertainment.

Our flagship apps — Retake AI, Cleanup, Learna, and DramaPops — and many of them lead their categories globally. In 2024, we became the most downloaded app publisher on iOS, driven by cutting‑edge AI research, sharp data‑driven execution, and a relentless focus on product and marketing.

We’re a team of 300+ people across İstanbul and Barcelona who bring curiosity, passion, trust, and ownership to everything we build. Recognized as a #1 LinkedIn Top Startup and a Great Place to Work in Europe, Codeway is where ambitious people do their life’s best work.

We’re building the next generation of consumer tech and reimagining what mobile apps can be.

This is Codeway. This is our way. Join us.

POSITION

We’re looking for a Senior Site Reliability Engineer to own and mature reliability, performance, and security across our growing platform. This role sits at the intersection of Engineering, Infrastructure, and Security, helping design, operate, and continuously improve the systems that keep dozens of consumer apps running for users around the world.

You’ll work closely with product engineering teams to make reliability measurable rather than assumed. That means defining and enforcing SLIs, SLOs, and error budgets; operating and hardening a multi‑cluster Kubernetes environment; building the observability that catches problems before users feel them; and leading incident response when things break. It’s a hands‑on role with real ownership over how reliability and security evolve as we scale.

Several parts of our reliability practice are still early in their maturity. We’re looking for someone who enjoys building the standards, processes, tooling, and automation that will form the foundation of our SRE function — not someone waiting for a playbook to already exist.

We welcome applicants from all backgrounds and experiences. If you’re excited about running systems at consumer scale and believe you could be a strong fit, we encourage you to apply, even if your experience doesn’t align perfectly with every qualification listed below.

WHAT YOU’LL BE DOING
Reliability, SLOs, Observability
  • Define, instrument, and report on SLIs, SLOs, and error budgets across critical services, so reliability decisions are driven by data rather than opinion.
  • Own observability end‑to‑end — metrics, logs, traces, dashboards, and alerting — and drive measurable reductions in detection and resolution times.
  • Reduce alert noise and false positives so on‑call engineers can trust what wakes them up.
  • Run reliability reviews and an error‑budget policy that shapes how teams prioritize between shipping and stability.
Kubernetes & Platform Operations
  • Operate, scale, and upgrade our multi‑cluster Kubernetes (GKE) environment: cluster lifecycle, autoscaling, networking, ingress, and resource management.
  • Act as the deep‑expert escalation point for cluster and platform issues across dozens of services.
  • Own capacity planning, performance, and cloud cost efficiency, balancing spend against reliability targets.
  • Build self‑service platform tooling that lets product teams move quickly without needing to become infrastructure experts.
Security & Resilience
  • Embed security into the platform through RBAC and least‑privilege, secrets management, image and dependency scanning, network policies, and a disciplined patching cadence.
  • Partner with the security function on vulnerability remediation, audit readiness, and secure‑by‑default infrastructure.
  • Own disaster recovery: define and regularly validate RTO/RPO targets through DR drills and failure testing.
  • Contribute to architecture and production‑readiness reviews so reliability and security are designed in, not bolted on.
Incident Response & Automation
  • Lead the on‑call rotation and act as incident commander during production incidents.
  • Run blameless postmortems, quantify impact, and track corrective actions through to closure so the same failure doesn’t recur.
  • Build and maintain Infrastructure as Code (Terraform) and CI/CD pipelines, enforcing GitOps and progressive delivery with automated rollbacks.
  • Systematically identify, measure, and eliminate operational toil through automation, protecting engineering time for high‑leverage work.
WHAT YOU’LL BRING?
  • Experience operating high‑traffic, always‑on production systems at meaningful scale, typically gained over 5–8 years in SRE, Platform, or DevOps roles.
  • Hands‑on production Kubernetes experience — you’ve run clusters day to day, through upgrades, autoscaling, and real troubleshooting under load, not just deployed to them.
  • A strong cloud engineering background, along with solid Linux and networking fundamentals.
  • A track record of defining and operating with SLOs and error budgets, and comfort being measured on reliability outcomes.
  • Experience with Infrastructure as Code and CI/CD pipeline design — you treat infrastructure and delivery as code.
  • Depth in observability tooling: instrumentation, dashboarding, and alert design.
  • A genuine security‑first mindset, where least‑privilege, secrets hygiene, and vulnerability management are habits rather than afterthoughts.
  • Scripting and automation fluency in at least one language, used to build tooling and remove toil.
  • Incident‑command experience: owning on‑call, running blameless postmortems, and driving resolution times down over time.
  • Ability to communicate clearly with both engineers and leadership, especially under pressure.
NICE TO HAVE
  • Experience with high‑scale consumer or mobile app backends, or with AI/ML inference workloads and their scaling characteristics.
  • Experience with GitOps and progressive‑delivery patterns such as canary and blue‑green rollouts.
  • Familiarity with service mesh, API gateways, or multi‑region and multi‑cluster topologies.
  • Cloud cost optimization and FinOps discipline at scale.
  • Exposure to compliance initiatives (SOC 2, ISO 27001, GDPR) and broader DevSecOps practice.
  • Chaos engineering or resilience testing experience.
  • Relevant certifications in Kubernetes, cloud, or DevOps disciplines.
  • Experience supporting many independent services and teams concurrently in a fast‑shipping, product‑led environment.
OUR ENVIRONMENT

You’ll help operate and improve a modern, cloud‑native environment built around:

  • Kubernetes (GKE) and containerized workloads
  • Google Cloud Platform (GCP)
  • Terraform and Infrastructure as Code
  • CI/CD and GitOps tooling
  • Modern observability (metrics, logs, traces, alerting)
  • Cloudflare CDN and edge

Experience with these exact platforms is beneficial but not required. We value strong fundamentals, curiosity, and the ability to quickly learn new technologies and environments.

WHAT SUCCESS LOOKS LIKE

Within your first 12 months, you’ll help establish and mature key reliability capabilities, including:

  • Clear SLIs, SLOs, and error budgets live and reported for our most critical services.
  • A measurable reduction in detection and resolution times
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer (Databases SRE)
Staff Software Engineer (Databases SRE)

Grafana • Banyoles

Remote
EUR 120,000 - 180,000
30 days vacation
Healthcare
Pension plan
+6
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Ribarroja del Turia

On-site
EUR 40,000 - 70,000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps
Competitive compensation: USD-based pay with education, fitness, and team activity budgets
Exciting projects: Modern solutions with Fortune 500 and top product companies
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Madrid

On-site
EUR 45,000 - 60,000
Professional growth
Competitive compensation
Exciting projects
+1
Site Reliability Engineer
Site Reliability Engineer

Emburse, Inc. • Barcelona

On-site
EUR 90,000 - 130,000
Flexible spending accounts
Generous paid time off
Paid parental leave
+9
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Tamarind Intelligence • Barcelona

On-site
EUR 55,000 - 70,000
Hybrid work model
Salary 55k-70k€ annually
Comprehensive health insurance
Senior SRE Engineer
Senior SRE Engineer

Lever, Inc. • Spain

Remote
EUR 90,000 - 130,000
Fully remote work
Medical insurance
Learning & development budget
+3
Senior SAP Expert
Senior SAP Expert

3530 Kyndryl España, S.A.U. • Madrid

Hybrid
EUR 90,000 - 130,000
Senior DevOps Engineer (Cloud-Native, AI-Driven Platform)
Senior DevOps Engineer (Cloud-Native, AI-Driven Platform)

SmartRecruiters, Inc. • Benidoleig

On-site
EUR 70,000 - 95,000
Senior DevOps Engineer (Cloud-Native, AI-Driven Platform)
Senior DevOps Engineer (Cloud-Native, AI-Driven Platform)

SmartRecruiters, Inc. • Sasamón

On-site
EUR 70,000 - 100,000
Senior DevOps Engineer
Senior DevOps Engineer

Rippling, Inc. • Bilbao

On-site
EUR 70,000 - 95,000
Medical, dental, vision insurance
Employee referral bonus
Wellness programs
+1