Senior Site Reliability Engineer

Randstad (Schweiz) AG

Madrid

Presencial

EUR 70.000 - 100.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Health Insurance
Paid Time Off (PTO)
Paid Holidays
Remote Work
Professional Development

Descripción de la vacante

SCALIS is seeking a Senior Site Reliability Engineer to own the reliability, scalability, and operational excellence of our platform in Madrid. You will partner with Engineering, Security, and Product to build resilient infrastructure, raise observability, and strengthen incident response.

You will design Kubernetes-based infrastructure on AWS, manage CI/CD and GitOps workflows (ArgoCD), implement autoscaling, and contribute to IaC with Terraform while exploring AI tooling to accelerate

Formación

  • 5+ years of experience in site reliability engineering.
  • Experience owning reliability for production systems — defining SLOs or running error budgets.
  • Deep hands-on experience with Kubernetes/Helm on EKS in production.
  • Experience with Cloudflare (CDN, WAF, etc).
  • Experience with CI/CD or general GitOps deployment pattern experience.
  • Has led or significantly contributed to incident response and postmortem processes.
  • Working knowledge of AWS core services (networking, IAM, compute) beyond just EKS.

Responsabilidades

  • Own and improve platform reliability (SLOs/SLIs), capacity planning, and production readiness
  • Design, build, and maintain Kubernetes-based infrastructure and deployment workflows
  • Operate and evolve our AWS footprint with a security-first mindset
  • Improve CD/GitOps practices (ArgoCD) and deployment safety
  • Build autoscaling strategies for services and workloads
  • Lead incident response: on-call, triage, mitigation, post-mortems, and preventative follow-through
  • Strengthen observability across services: metrics, logs, traces, and alerting
  • Partner with application teams to tune performance, reduce toil, and improve operational maturity
  • Improve infrastructure-as-code practices and maintain Terraform modules
  • Contribute to evaluating and integrating AI tooling and MCP tools to accelerate operational workflows

Conocimientos

SRE experience
SLOs/SLIs
Kubernetes
Helm
EKS
AWS
CI/CD
GitOps
Incident response
Observability

Herramientas

ArgoCD
Terraform
Cloudflare
OpenTelemetry

Descripción del empleo

Site Reliability Engineer (SRE)

You’ll own the reliability, scalability, and operational excellence of the systems that power our platform. You’ll partner closely with Engineering, Security, and Product to build resilient infrastructure, improve developer experience, and raise the bar on observability and incident response.

What you’ll be doing:
  • Own and improve platform reliability (SLOs/SLIs), capacity planning, and production readiness
  • Design, build, and maintain Kubernetes‑based infrastructure and deployment workflows
  • Operate and evolve our AWS footprint (networking, compute, storage, IAM) with a security‑first mindset
  • Improve our CD/GitOps practices (ArgoCD) and deployment safety (progressive delivery, rollbacks, guardrails)
  • Build autoscaling strategies for services and workloads (KEDA where appropriate)
  • Lead incident response: on‑call, triage, mitigation, post‑mortems, and preventative follow‑through
  • Strengthen observability across services: metrics, logs, traces, and alerting (OpenTelemetry + dashboards)
  • Partner with application teams to tune performance, reduce toil, and improve operational maturity
  • Improve infrastructure‑as‑code practices and maintain Terraform modules and environments
  • Contribute to evaluating and integrating AI tooling and MCP tools to accelerate operational workflows
What you should bring:
  • 5+ years of experience in site reliability engineering
  • Experience owning reliability for production systems — defining SLOs or running error budgets
  • Deep hands‑on experience with Kubernetes/Helm on EKS in production
  • Experience with Cloudflare (CDN, WAF, etc)
  • Experience with CI/CD or general GitOps deployment pattern experience
  • Has led or significantly contributed to incident response and postmortem processes
  • Working knowledge of AWS core services (networking, IAM, compute) beyond just EKS
Nice to have:
  • Experience evaluating or building AI‑assisted ops tooling (MCP, agentic runbooks)
  • OpenTelemetry / observability pipeline design
  • Experience with Terraform
Benefits:
  • Health Insurance
  • Paid Time Off (PTO)
  • Paid Holidays
  • Remote Work
  • Professional Development
EEO Statement:

Fountain is a proud equal opportunity workplace. We welcome applicants of any educational background, gender identity and expression, sexual orientation, religion, ethnicity, age, socioeconomic status, disability, and veteran status.

By submitting an application, you confirm that you have read our Privacy Policy and agree that we may process and retain your personal data for the purpose of recruitment in accordance with applicable data protection laws.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Fountain • Madrid

Presencial
EUR 70.000 - 110.000
Competitive health plans
Retirement plan
Flexible vacation policy
+2
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Madrid

Híbrido
EUR 45.000 - 60.000
Professional growth
Competitive compensation
Exciting projects
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Ribarroja del Turia

Híbrido
EUR 40.000 - 70.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps
Competitive compensation: USD-based pay with education, fitness, and team activity budgets
Exciting projects: Modern solutions with Fortune 500 and top product companies
+1
Senior SRE Engineer
Senior SRE Engineer

Codeway • Barcelona

Presencial
EUR 85.000 - 120.000
Senior SRE — Global Remote, Scalable Systems
Senior SRE — Global Remote, Scalable Systems

Fountain • Madrid

Presencial
EUR 70.000 - 110.000
Competitive health plans
Retirement plan
Flexible vacation policy
+2
Site Reliability Engineer ID62591
Site Reliability Engineer ID62591

AgileEngine • Ribarroja del Turia

Híbrido
EUR 40.000 - 60.000
Mentorship
TechTalks
USD-based pay
+2
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Hydrolix • España

Presencial
EUR 70.000 - 90.000
Senior Site Reliability Engineer — Reliability & Platform Ops
Senior Site Reliability Engineer — Reliability & Platform Ops

Codeway • Barcelona

Presencial
EUR 85.000 - 120.000
Senior Site Reliability Engineer (Guardicore AI Platform) - Remote
Senior Site Reliability Engineer (Guardicore AI Platform) - Remote

Akamai Technologies • Madrid

Presencial
EUR 70.000 - 110.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

F. Hoffmann-La Roche AG • Sant Cugat del Vallès

Presencial
EUR 50.000 - 75.000