Staff Site Reliability Engineer

levistraussandco

España

Presencial

EUR 120.000 - 180.000

Jornada completa

Hace 6 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Descripción de la vacante

Levi Strauss & Co. in Spain seeks an exceptional Staff Site Reliability Engineer to own the reliability, scalability, and operability of our data and AI platforms.

You will lead with GCP expertise across multi-cloud environments to support the design of our iconic jeans and global operations. As a hands-on technical leader, you will reduce toil, improve observability, and define SLOs/SLIs, while mentoring junior engineers and partnering with Data, Security, and Product teams to embed reliability

Formación

  • Master's degree in Computer Science, Engineering, or related field (or equivalent practical experience).
  • 10+ years in SRE, DevOps, or Platform Engineering in large-scale production.
  • Deep hands-on GCP experience (GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, Vertex AI).
  • Terraform and Helm with GitOps (ArgoCD, Flux).
  • Strong observability skills: tracing, logging, metrics, alerting (Cloud Monitoring, Datadog, Prometheus/Grafana).
  • Experience defining/operating SLOs/SLIs/error budgets in production.
  • Data security: IAM, encryption, secrets, network policies, compliance.
  • Experience in multi-cloud (GCP + Azure).
  • Ability to lead without authority.

Responsabilidades

  • Define, instrument, and enforce SLOs, SLIs across platform services.
  • Reduce MTTD/MTTR via observability, alerting, and runbook incident response.
  • Lead blameless post-mortems and drive reliability improvements.
  • Identify and reduce operational toil; maintain guardrails to keep toil under 50%.
  • Build self-serve infrastructure for teams to provision and operate resources.
  • Automate deployment pipelines and IaC workflows (Terraform, Helm, GitOps).
  • Architect workloads across GCP (GKE, Cloud Run, BigQuery, Pub/Sub, etc.) and Azure.
  • Design self-healing infra, auto-scaling, and capacity planning for high availability.
  • Champion data security and governance: encryption, IAM, secrets, audit logs.
  • Apply SRE to agentic AI workloads; define reliability for LLM/multi-agent systems.
  • Partner with AI Platform teams to productionize pipelines with monitoring and drift detection.
  • Drive adoption of AI-assisted operations tooling for observability and incident management.
  • Lead and mentor junior/mid-level SREs; collaborate with cross-functional teams.
  • Communicate platform health and reliability roadmap to technical and exec audiences.

Conocimientos

SRE experience
DevOps
Platform engineering
Leadership & mentoring
Communication with executives

Educación

Master's degree in Computer Science or related field

Herramientas

Terraform
Helm
GitOps (ArgoCD, Flux)
Cloud Monitoring
Datadog
Prometheus/Grafana
ArgoCD
Flux

Descripción del empleo

Job Location: Spain

Calling all originals: At Levi Strauss & Co., you can be yourself - and be part of something bigger. We're a company of people who like to forge our own path and leave the world better than we found it. Who believe that what makes us different makes us stronger. So add your voice. Make an impact. Find your fit - and your future.

We're seeking an exceptional Staff Site Reliability Engineer to join our Data & AI Platform Engineering team. In this role, you'll own and elevate the reliability, scalability, and operability of our enterprise data and AI platforms - the platforms that power everything from the design of our iconic jeans to the optimization of our global retail and supply chain.

As a hands-on technical leader, you'll embody the principles of Google's SRE discipline: eliminating toil, engineering for reliability, and building a culture of shared ownership between development and operations. This is a unique opportunity to shape how a legendary brand runs production at scale on Google Cloud Platform, with a growing multi-cloud footprint across GCP and Azure.

About the Job
Reliability & Incident Management
  • Define, instrument, and enforce SLOs, SLIs, and error budgets across all platform services, ensuring alignment with business and product commitments
  • Drive continuous reduction in MTTD and MTTR through improved observability, automated alerting, and runbook-driven incident response
  • Lead blameless post-mortems and translate findings into durable reliability improvements, ensuring systemic issues are eliminated rather than patched
Toil Reduction & Automation
  • Systematically identify, measure, and eliminate operational toil; track toil percentage per sprint and enforce guardrails to keep it below 50% of engineering capacity
  • Build and maintain self-serve infrastructure capabilities - enabling product and data engineering teams to provision, scale, and operate their own resources safely and consistently
  • Automate deployment pipelines, configuration management, and operational workflows using Infrastructure-as-Code principles (Terraform, Helm, GitOps)
Platform Engineering & Architecture
  • Serve as the primary GCP subject matter expert - architecting and optimizing workloads across GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI
  • Lead multi-cloud architecture decisions across GCP and Azure, ensuring consistent security posture, cost efficiency, and operational practices across environments
  • Design and implement self-healing infrastructure patterns , auto-scaling strategies, and capacity planning models to support high-availability data and AI platforms
  • Champion data security and governance best practices - including encryption at rest and in transit, IAM least-privilege, secrets management, and audit logging
AI, Agentic Systems & Modern Observability
  • Apply SRE principles to agentic AI workloads - defining reliability expectations for LLM-based and multi-agent systems, including latency SLOs, fallback patterns, and model observability
  • Partner with AI Platform teams to productionize agentic pipelines with robust monitoring, drift detection, and rollback capabilities
  • Drive adoption of AI-assisted operations tooling to enhance observability, anomaly detection, and predictive incident management
Leadership & Culture
  • Guide and mentor junior and mid-level SREs - conducting code reviews, running reliability reviews, and elevating the team's engineering craft
  • Collaborate cross-functionally with Data Engineering, Software Engineering, Security, and Product teams to embed reliability as a shared value from design through deployment
  • Champion a culture of psychological safety, continuous learning, and reliability excellence
  • Communicate platform health, risk posture, and reliability roadmaps clearly to both technical and executive audiences
About You
Required Qualifications
  • Master's degree in Computer Science, Engineering, or related field (or equivalent practical experience)
  • 10+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering, with a strong track record in large-scale production environments
  • Deep, hands-on expertise in GCP - including GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI
  • Proficiency with Infrastructure-as-Code tools (Terraform, Helm) and GitOps workflows (ArgoCD, Flux)
  • Strong command of observability tooling - distributed tracing, structured logging, metrics pipelines, and alerting platforms (e.g., Cloud Monitoring, Datadog, Prometheus/Grafana)
  • Proven experience defining and operating against SLOs, SLIs, and error budgets in production environments
  • Solid understanding of data security principles : IAM, encryption, secrets management, network policies, and compliance frameworks
  • Experience with multi-cloud environments (GCP + Azure), including cross-cloud networking, identity federation, and cost governance
  • Demonstrated ability to lead without authority - influencing engineers acr
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Stellar Cyber • Banyoles

Presencial
EUR 90.000 - 130.000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+6
Staff SRE, Data & AI Platform (GCP/Azure)
Staff SRE, Data & AI Platform (GCP/Azure)

levistraussandco • España

Presencial
EUR 120.000 - 180.000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Madrid

Presencial
EUR 45.000 - 60.000
Professional growth
Competitive compensation
Exciting projects
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Ribarroja del Turia

Presencial
EUR 40.000 - 70.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps
Competitive compensation: USD-based pay with education, fitness, and team activity budgets
Exciting projects: Modern solutions with Fortune 500 and top product companies
+1
Senior Site Reliability Engineer - Platform Reliability (Resilience)
Senior Site Reliability Engineer - Platform Reliability (Resilience)

Elastic • España

A distancia
EUR 76.000 - 102.000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Hydrolix • España

Presencial
EUR 70.000 - 90.000
Senior Software Engineer - Reliability, Infrastructure, and Tooling
Senior Software Engineer - Reliability, Infrastructure, and Tooling

Jobgether • España

A distancia
EUR 117.000 - 261.000
Equity participation
Fully remote
Health, dental, and vision
+1
Site Reliability Engineer
Site Reliability Engineer

Emburse • Barcelona

Presencial
EUR 90.000 - 130.000
Flexible spending accounts
Generous paid time off
Paid parental leave
+9
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Tamarind Intelligence • Barcelona

Presencial
EUR 55.000 - 70.000
Hybrid work model
Salary 55k-70k€ annually
Comprehensive health insurance
Principal Software Engineer
Principal Software Engineer

Love Brand Group • Valladolid

Presencial
EUR 110.000 - 160.000