Senior Site Reliability Engineer

Jobgether

Deutschland

Vor Ort

EUR 90.000 - 130.000

Vollzeit

Vor 6 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Fully remote
High ownership
Tech exposure

Zusammenfassung

Jobgether is seeking a Senior Site Reliability Engineer based in Germany to own cloud infrastructure and drive reliability, observability, and operational excellence. You will shape scalable infrastructure on Google Cloud, enable development teams with automation, and lead incident management and post-incident reviews.

Responsibilities include designing scalable GCP infra, building internal tooling, improving SLOs/SLIs, and optimizing cloud costs while collaborating with cross-functional teams.

Qualifikationen

  • 3+ years of professional SRE or production-focused experience.
  • Strong hands-on GCP expertise with cost optimization and governance.
  • Practical experience managing Kubernetes clusters in production.

Aufgaben

  • Design, build, and evolve scalable cloud infrastructure on GCP.
  • Develop automation and self-service tooling to boost team autonomy.
  • Advance observability with metrics, logging, tracing, and distributed systems.
  • Establish cost visibility and run cloud governance initiatives.

Kenntnisse

GCP
Kubernetes
Terraform
Python

Tools

Prometheus
Grafana
OpenTelemetry

Jobbeschreibung

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Germany. This role offers the opportunity to take ownership of cloud infrastructure and help shape reliability, observability, and operational excellence across a growing engineering organization. You will design and evolve scalable infrastructure on Google Cloud while enabling development teams through automation and self-service tooling. The role combines hands-on technical engineering with ownership of reliability practices, incident management, and operational improvement. You will strengthen observability capabilities and use data-driven metrics to improve system performance and software delivery. You will also help optimize cloud usage and costs while establishing sustainable infrastructure governance. This is an ideal opportunity for an experienced SRE who enjoys solving complex operational challenges and working closely with cross-functional engineering teams.

Accountabilities
  • Design, build, and continuously evolve scalable cloud infrastructure using Google Cloud Platform (GCP).
  • Develop internal tooling, automation, and self-service capabilities that improve engineering team autonomy and operational efficiency.
  • Advance observability across metrics, logging, tracing, and distributed systems to improve system visibility and reduce mean time to recovery.
  • Establish infrastructure cost visibility and lead governance and optimization initiatives to improve cloud efficiency.
  • Champion reliability engineering practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and DORA metrics.
  • Define and improve incident management processes, lead incident response, and facilitate post-incident reviews focused on identifying and addressing root causes.
  • Participate in an on-call rotation and help ensure reliable operation of production systems.
  • Partner with engineering and product teams to identify reliability risks, improve operational practices, and support the delivery of resilient software.
  • Contribute to continuous improvements across infrastructure, deployment, monitoring, and operational workflows.
Requirements
  • 3+ years of professional experience in Site Reliability Engineering or production-focused SRE roles.
  • Strong hands-on proficiency with Google Cloud Platform (GCP), including cloud cost optimization and governance.
  • Practical experience managing Kubernetes clusters and workloads in production environments.
  • Infrastructure as Code experience using Terraform, Deployment Manager, or comparable tools.
  • Strong scripting and automation capabilities using Python, Bash, Go, or similar languages.
  • Experience building and maintaining observability solutions using tools such as Prometheus, Grafana, and OpenTelemetry, along with logging and distributed tracing.
  • Proven experience with incident management, post-incident reviews, production troubleshooting, and on-call responsibilities.
  • Ability to define, implement, and operationalize SLOs, SLIs, and error budgets.
  • Understanding of DORA metrics and their application to software engineering and delivery workflows is an advantage.
  • Experience designing and evolving cloud infrastructure at scale is highly desirable.
  • Background working in AI/ML, geospatial technology, or similarly data-intensive environments is a plus.
  • Strong problem-solving, communication, collaboration, and ownership skills.
  • Ability to work effectively in a fully distributed and cross-functional engineering environment.
  • Candidates must be based in Canada or another eligible location within North America, the EU, or the UK.
  • Visa sponsorship is not available for this role.
Benefits
  • Fully remote work with flexibility across eligible locations.
  • Opportunity to work on challenging cloud infrastructure, reliability, and observability initiatives.
  • High degree of ownership and influence over infrastructure and operational engineering practices.
  • Opportunity to build automation and self-service tooling that directly improves engineering productivity.
  • Exposure to modern cloud-native technologies, including GCP, Kubernetes, Terraform, Prometheus, Grafana, and OpenTelemetry.
  • Opportunity to contribute to reliability engineering practices involving SLOs, SLIs, error budgets, incident management, and DORA metrics.
  • Salary discussed during the interview process and determined based on experience and location.
  • Opportunity to collaborate with experienced engineering teams in an innovative AI/ML and technology-driven environment.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Site Reliability Engineering Architect
Site Reliability Engineering Architect

Cavendish Professionals • Berlin

Hybrid
EUR 80.000 - 110.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Visa Hunt • Deutschland

Vor Ort
EUR 90.000 - 150.000
Home office budget
Learning & development budget of €1000
Competitive salary
+5
Senior Engineer (Platform)
Senior Engineer (Platform)

Jobgether • Deutschland

Vor Ort
EUR 120.000 - 160.000
Senior DevOps (GPC)
Senior DevOps (GPC)

Aether Biomedical • Deutschland

Hybrid
EUR 100.000 - 150.000
Health insurance
Life insurance
Sick leave days
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

B Capital • Deutschland

Remote
EUR 46.000 - 105.000
Work from anywhere
Flexible paid time off
Mental health support services
+3
(Senior) Site Reliability Engineer (m/f/d) in Berlin or Konstanz
(Senior) Site Reliability Engineer (m/f/d) in Berlin or Konstanz

United States Digital Space LLC • Deutschland

Hybrid
EUR 60.000 - 90.000
Hybrid working arrangements
Flexible hours
Subsidized sports or yoga courses
+1
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

FACT-Finder • Pforzheim

Hybrid
EUR 90.000 - 140.000
Hybrid work model
Impact on product reliability
Competitive compensation
+1
Senior Site Reliability Engineer (m/w/d)
Senior Site Reliability Engineer (m/w/d)

Impower • München

Hybrid
EUR 70.000 - 90.000
Flexible hours
Ownership in projects
Diverse team culture
Google Cloud Architect
Google Cloud Architect

Franklin Fitch • Berlin

Hybrid
EUR 99.000 - 121.000
Salary up to €110,000
30 days holiday
25 days workation in EU
+1
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • Stuttgart

Vor Ort
EUR 70.000 - 100.000