Staff Software Engineer (Databases SRE)

Grafana

Banyoles

A distancia

EUR 120.000 - 180.000

Jornada completa

14 días+
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca la empresa.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

30 days vacation
Healthcare
Pension plan
Learning and development stipend
Remote-friendly
Co-working space stipend
Headspace mindfulness
Employee assistance program
Parental leave

Descripción de la vacante

Grafana Labs in Spain is seeking a Staff Software Engineer - SRE to own production reliability for Grafana Cloud databases. You will design automation, improve observability, drive per-tenant SLOs, and lead incident response for high-SLA environments, partnering with product engineering to scale reliability across AWS, GCP, and Azure.

You’ll contribute to design reviews, mentor others, and help reduce toil while delivering scalable, fault-tolerant systems.

Formación

  • 8+ years of engineering experience with 4+ in SRE/production engineering.
  • Strong knowledge of Kubernetes in cloud environments (AWS/GCP/Azure).
  • Experience designing SLOs and incident response processes.
  • Proficient in at least one language (Go, Python, Java).

Responsabilidades

  • Own production reliability for high-SLA and complex customer environments.
  • Lead customer-impacting incident response and post-incident reviews.
  • Design and implement automation to scale reliability practices.
  • Partner with product engineering squads for reliability.
  • Improve observability and alerting to reduce noisy escalations.
  • Influence feature design for production scalability and operability.

Conocimientos

SRE/Production engineering
Kubernetes on cloud
Go/Python/Java
Linux internals
Incident response
Observability
Multi-tenant systems

Herramientas

Helm
Terraform
Jsonnet
CI/CD

Descripción del empleo


  • We are looking for a Staff Software Engineer - SRE to help us support our highest value Grafana Cloud customers by increasing the reliability of our Cloud databases that are based on Mimir, Loki, Tempo, and Pyroscope. We provide these databases as a SaaS product from AWS, GCP, and Azure across all regions

  • The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on ensuring that Grafana Cloud’s database products deliver exceptional reliability for our highest-SLA customers. In this role, you will:

  • Partner closely with product engineering squads (embedded model)

  • Own production reliability for high-SLA and complex customer environments

  • Design and implement automation to scale our reliability practices

  • Ensuring our customers meet our SLO targets

  • Define and evolve per-tenant SLOs and reliability models

  • Proactively reduce SLO burn to prevent repeat incidents

  • Serving as a primary escalation point and on-call for relevant incidents

  • Lead customer-impacting incident response and post-incident reviews

  • Contribute to design docs and code reviews

  • Influence feature design to ensure production scalability and operability

  • Build automation to eliminate toil where needed

  • Improve alert quality and reduce noisy escalations

  • We seek a staff software engineer operating at the intersection of customer needs, production systems, and product engineering

  • Regular 1:1s to with your manager and colleagues

  • Reviewing and creating SLOs, proactively investigating ways in which we can further reduce budget burn for those SLOs, which can be self-directed or as the result of learnings from incidents, and may include improvements to monitoring, automation, increasing self-healing, auto-scaling, etc

  • Improve observability of customers within their environments

  • Designing and implementing solutions to ensure reliability and scalability of our environments can meet rapidly increasing demands

  • Develop fault-tolerant design patterns ensuring that we are considering reliability at all stages of the service lifecycle

  • Collaborating with our Engineering Leaders to help define and influence product strategy, roadmaps and technical designs

  • Participate in PR review and collaborating with other engineers on their Design Docs

  • Teach others about Site Reliability Engineering and communicate best practices to be applied early in development of new features and functionality

  • Participate in Incident Response when applicable, including investigation through to resolution, PIR, and communication with customers via Bridge calls where necessary


Benefits


  • Vacation: Balance is key. Our team enjoys 30 days of paid vacation each year on top of national holidays, parental leave, and sick leave. We also take a breather on a number of Grafana Shutdown Days each year

  • Healthcare: We’re proud to provide health coverage or stipends for our colleagues in the US, UK, Canada, the Netherlands, Sweden, Singapore, and India

  • Retirement planning: There’s no time like the present to start saving for your future. We make employer contributions into the pension pots of our team members in the US, UK, Canada, the Netherlands, Sweden, and Germany

  • Professional development: On top of a $1,500 annual learning and development stipend, Grafanistas have thousands of on-demand courses at their fingertips to help them grow professionally. Want to attend a conference or training? Go ahead. Just pass on what you learned

  • Work location: Vast majority of our roles are fully remote, focused on hiring the best talent and allowing you to perform from the comfort of your home. If you fancy a change of scene, we’ll also reimburse you up to $175 a month for a personal co-working space

  • Choice of tech: There’s no one-size‑fits‑all when it comes to the tech required to do your job. Choose the laptop and accessories you need when you join us, and we’ll refresh them every three years

  • Mindfulness: When you join the team, you can sign up for a complimentary subscription to Headspace to take advantage of the benefits of mindfulness and meditation. Our wellbeing resource group also organize sessions run by fellow Grafanistas or external trainers

  • Global Employee Assistance Program: We offer all team members a 100% confidential support service with 24/7 365 access to professionally qualified counsellors and specialists

  • Paid parental leave: Grafana offers paid parental leave to all eligible new parents. This offers Grafanistas time to bond with and care for their children in the first year after birth or adoption



  • You may not meet every requirement, and that’s okay.

  • If this role excites you, we’d love you to raise your hand for what could be a truly career-defining opportunity

  • Ability to reason about performance, scaling, and failure modes

  • Experience operating multi-tenant systems in production

  • Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling

  • Strong experience designing and implementing SLOs

  • Ability to partner deeply with product engineering teams

  • Experience with one or more programming languages (e.g. Go, Python, Java, etc)

  • Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.)

  • 8+ years engineering experience, 4+ in SRE/CRE/production engineering. Strong preference for those with formal customer reliability engineering experience

  • Excellent problem-solving and troubleshooting skills

  • Experience with calmly and actively participating in blame‑free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post‑mortem documents)

  • Strong experience with technical leadership, leading a team through projects, mentoring other engineers on the team and serving as a force‑multiplier

  • Comfortable working within an engineering team where individuals are encouraged to have a strong sense of autonomy and self‑direction

  • We highly value those who are intellectually curious, who default to transparency, possess a high bias towards action, and who are also kind (this is important!)

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Backend Engineer (Databases, Analytics)
Senior Backend Engineer (Databases, Analytics)

Grafana • Banyoles

Presencial
EUR 90.000 - 130.000
30 days vacation holidays
Healthcare stipend
Pension contributions
+5
Senior Backend Engineer - Databases - Analytics | Spain | Remote
Senior Backend Engineer - Databases - Analytics | Spain | Remote

Grafana • España

A distancia
EUR 83.000 - 104.000
Remote work
Senior Backend Engineer - Databases Pyroscope | Spain | Remote New Spain (Remote)
Senior Backend Engineer - Databases Pyroscope | Spain | Remote New Spain (Remote)

Grafana • España

A distancia
EUR 82.988 - 99.586
RSUs
Remote-first culture
Staff Backend Engineer - Grafana Second Horizon | Spain | Remote
Staff Backend Engineer - Grafana Second Horizon | Spain | Remote

Grafana • Madrid

Presencial
EUR 94.000 - 113.000
Remote-first team
RSUs
30 days annual leave
+1
Senior Backend Engineer - Databases - Analytics | UK | Remote
Senior Backend Engineer - Databases - Analytics | UK | Remote

Showcify, Inc. • España

A distancia
GBP 91.000 - 115.000
Equity
Bonus (if applicable)
Senior Backend Engineer - Databases - Analytics | UK | Remote
Senior Backend Engineer - Databases - Analytics | UK | Remote

Grafanalabs • España

A distancia
GBP 91.000 - 115.000
Equity
Bonus
30 days annual leave
+1
Site Reliability Engineer
Site Reliability Engineer

Emburse, Inc. • Barcelona

Presencial
EUR 90.000 - 130.000
Flexible spending accounts
Generous paid time off
Paid parental leave
+9
Staff SRE, Cloud Databases | SLOs, Multi-Tenant, Kubernetes
Staff SRE, Cloud Databases | SLOs, Multi-Tenant, Kubernetes

Grafanalabs • España

A distancia
EUR 94.025 - 112.830
Engineering Manager - Observability | Spain | Remote
Engineering Manager - Observability | Spain | Remote

Grafana • España

A distancia
EUR 94.000 - 117.000
RSUs
30 days annual leave
Grafana Shutdown Days
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Ribarroja del Turia

Presencial
EUR 40.000 - 70.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps
Competitive compensation: USD-based pay with education, fitness, and team activity budgets
Exciting projects: Modern solutions with Fortune 500 and top product companies
+1