Senior Cloud SRE — Kubernetes, AWS & Reliability Lead

Nexthink

Madrid

Híbrido

EUR 60.000 - 86.000

Jornada completa

hace 40 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Destaca para este puesto: genera un currículum y una carta de presentación adaptados en cuestión de un minuto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Hybrid work model
Unlimited vacation
Volunteer days
Relocation package
Office with lakeview

Descripción de la vacante

Nexthink is seeking a Senior Site Reliability Engineer to join our global SRE team. You will strengthen our cloud-native infrastructure, deploy and monitor multi-tenant SaaS platforms, and collaborate with 50+ Product Engineering teams to ensure reliability, performance, and security across all environments.

The role emphasizes designing for resilience, implementing automation, and leading incident responses in a fast-paced, hybrid work setting.

Formación

  • 5+ years of experience as a Site Reliability Engineer or Platform Engineer with strong knowledge of software development best practices.
  • Strong hands-on experience with public cloud services (AWS, GCP, Azure) and supporting SaaS product.
  • Strong programming or scripting skills (e.g., Python, Go, Bash...), and experience with infrastructure-as-code (e.g. Terraform).
  • Proficiency with Kubernetes, container-based deployment (e.g., Docker) and related ecosystems (e.g., Helm).
  • Experience supporting multi-tenant microservices architectures.
  • Experience with CI/CD pipelines & tools (e.g., Jenkins, GitHub Actions, GitLab CI, FluxCD, Crossplane).
  • Experience with managing monitoring solutions (e.g. Datadog).
  • Comfortable participating in a rotating on-call schedule, managing critical incidents, and leading post-incident reviews.
  • At ease with operating and managing production systems, striking the right balance between urgency and methodology.
  • Strong system-level troubleshooting skills and a proactive mindset toward incident prevention.
  • Deep understanding of Linux systems, networking, and common troubleshooting practices.
  • Solid understanding of the network stack (e.g., TCP/IP, VPN, etc.), cloud architectures (VPC, subnets, firewalls, load balancers), service mesh (e.g., Istio) and storage (e.g., S3, EBS, etc).
  • Knowledge of zero-downtime deployment strategies, blue/green and canary releases.
  • Exposure to compliance standards such as SOC 2, ISO 27001, or HIPAA. FedRAMP experience is a big plus.
  • Experience with chaos engineering or resilience testing practices.
  • Excellent problem-solving skills, collaborative mindset, and a strong grasp of agile, iterative development.
  • Self-driven, highly organised, and capable of independently managing priorities.
  • Curiosity to learn new things and discover new technologies.
  • Strong communication, presentation, and team collaboration skills.
  • Excellent written and verbal skills in English.

Responsabilidades

  • Implement and manage cloud-native systems (AWS) using best-in-class tools and automation.
  • Operate and enhance Kubernetes clusters, deployment pipelines, and service meshes to support rapid delivery cycles.
  • Design, build, and maintain the infrastructure powering our multi-tenant SaaS platform with reliability, security, and scalability in mind.
  • Define and maintain SLOs, SLAs, and error budgets, and proactively address availability and performance issues.
  • Develop infrastructure-as-code (Terraform or similar) for repeatable and auditable provisioning.
  • Build internal platform tools and automation to support provisioning, monitoring, and operational efficiency.
  • Monitor infrastructure and applications ensuring high-quality user experiences.
  • Participate in a shared on-call rotation, responding to incidents, troubleshooting outages, and driving timely resolution and communication.
  • Act as an Incident Commander during the on-call duty and coordinate cross-team responses effectively to maintain an SLA.
  • Drive and refine incident response processes, reducing Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR).
  • Diagnose and resolve complex issues independently, minimizing the need for external escalation.
  • Work closely with software engineers to embed observability, fault tolerance, and reliability principles into service design.
  • Automate runbooks, health checks, and alerting to support reliable operations with minimal manual intervention.
  • Support automated testing, canary deployments, and rollback strategies to ensure safe, fast, and reliable releases.
  • Contribute to security best practices, compliance automation, and cost optimization.

Conocimientos

5+ years SRE
Public cloud (AWS/GCP/Azure)
Python/Go/Bash
Kubernetes/Docker/Helm
Multi-tenant microservices
CI/CD pipelines
Monitoring (Datadog)
On-call
Production operations
Troubleshooting
Linux/Networking
Canary/Blue-Green
Compliance (SOC2/ISO27001/HIPAA)
Chaos engineering
Agile
Self-motivation

Educación

Bachelor’s degree in Computer Science or equivalent

Herramientas

Terraform

Descripción del empleo

Nexthink is seeking a Senior Site Reliability Engineer to join our global SRE team. You will strengthen our cloud-native infrastructure, deploy and monitor multi-tenant SaaS platforms, and collaborate with 50+ Product Engineering teams to ensure reliability, performance, and security across all environments.

The role emphasizes designing for resilience, implementing automation, and leading incident responses in a fast-paced, hybrid work setting.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Hybrid Cloud-Native SRE: Kubernetes & Automation
Hybrid Cloud-Native SRE: Kubernetes & Automation

Nexthink SA • Madrid

Híbrido
EUR 60.000 - 80.000
Health insurance
Unlimited vacation
Flexible hours
+1
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Stellar Cyber • Banyoles

Presencial
EUR 90.000 - 130.000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+6
Senior Software & DevOps Engineer — Cloud & SRE (Hybrid)
Senior Software & DevOps Engineer — Cloud & SRE (Hybrid)

Nexthink • Madrid

Híbrido
EUR 60.000 - 80.000
Private Health Insurance
Meal vouchers
Unlimited vacation
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Nexthink SA • Madrid

Híbrido
EUR 60.000 - 80.000
Health insurance
Unlimited vacation
Flexible hours
+1
Senior SRE - Multi-Cloud Reliability (Remote)
Senior SRE - Multi-Cloud Reliability (Remote)

Palo Alto Networks • Madrid

Presencial
EUR 75.000 - 110.000
Senior SRE Engineer: Cloud-Native Automation Leader
Senior SRE Engineer: Cloud-Native Automation Leader

Stellar Cyber • España

A distancia
EUR 100.000 - 130.000
Senior SRE: Cloud, Kubernetes & Observability
Senior SRE: Cloud, Kubernetes & Observability

Stellar Cyber • España

Presencial
EUR 90.000 - 130.000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+6
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Nexthink • Madrid

Híbrido
EUR 60.000 - 86.000
Hybrid work model
Unlimited vacation
Volunteer days
+2
Senior Cloud SRE | AWS, AI-Driven Reliability | Hybrid
Senior Cloud SRE | AWS, AI-Driven Reliability | Hybrid

Comply365 • Barcelona

Híbrido
EUR 100.000 - 120.000
Laptop & monitor
Annual learning budget
Conference budget
+2
Senior Site Reliability Engineer - Cloud & DevOps
Senior Site Reliability Engineer - Cloud & DevOps

ThunderSoft • Madrid

Presencial
EUR 52.000 - 76.000