Senior Site Reliability Engineer / Cloud Operations Engineer (m/f/d)

RxREVU, Inc.

Berlin

Vor Ort

EUR 90.000 - 130.000

Vollzeit

Vor 8 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Thales in Germany seeks a Site Reliability Engineer to operate sovereign cloud services with 99.99%+ availability in distributed environments. You will monitor, troubleshoot, and resolve incidents, join a 24/7 on-call rotation, and collaborate with international teams to improve reliability and long-term solutions.

You will gain hands-on experience with Google Cloud technologies (Borg, Colossus, Spanner) and contribute to standardized incident response, playbooks, and automation to meet

Qualifikationen

  • Several years of experience in SRE, Cloud Operations, or related roles.
  • Experience maintaining business-critical production systems with high uptime.
  • Strong troubleshooting and incident management skills.
  • Familiarity with observability, monitoring, alerting, and RCA.
  • Proficiency in German and English.

Aufgaben

  • Operate and maintain mission-critical sovereign cloud services with high availability.
  • Monitor service health, SLIs, and SLOs across distributed systems.
  • Investigate, troubleshoot, and resolve complex production incidents.
  • Participate in a structured 24/7 on-call rotation (approx. one week every six weeks).
  • Collaborate with international teams to mitigate incidents and implement long-term improvements.
  • Build understanding of Google Cloud technologies and core GCP components.
  • Create and maintain technical documentation and standardize incident response procedures.
  • Lead post-incident reviews and root cause analyses, implementing preventive measures.
  • Identify opportunities for automation and improve operational efficiency.
  • Support secure cloud environments meeting regulatory requirements.
  • Contribute to reliability and security improvements across platforms.

Kenntnisse

SRE
Cloud operations
DevOps
Platform engineering
Infrastructure engineering
Production support
Incident management
Automation & IaC
German language
English language

Tools

GCP
Borg
Colossus
Spanner

Jobbeschreibung

Location: Berlin, Germany

We Say HI*
Site Reliability Engineer / Cloud Operations Engineer (f/m/d)

German companies and public administrations in this country are ready to accelerate their digital transformation and the use of AI-but they will never compromise on the security of their most sensitive data. This is where Thales in Germany, in partnership with Google Cloud and our new company currently being established, comes into play. With a new, 100% German business unit, we are providing a concrete response to the strict requirements of the BSI. What we are creating is a locally and fully autonomously operated "Trusted Cloud". It provides access to the broadest service portfolio on the market, while everything remains strictly under European jurisdiction. By combining German and French standards such as SecNumCloud, C5 and C3-A, we offer our customers unequaled resilience and business continuity. This is a turning point for our industry and a decisive step towards a strong, sovereign digital Europe.

Your mission as Site Reliability Engineer:
  • Operate and maintain mission-critical sovereign cloud services with availability targets of 99.99% and above.
  • Monitor service health, reliability, scalability, latency, and performance using Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
  • Investigate, troubleshoot, and resolve complex production incidents across large-scale distributed cloud environments.
  • Participate in a structured 24/7 on-call rotation (approximately one week every six weeks) to ensure continuous service availability.
  • Collaborate with Site Reliability Engineers, Cloud Infrastructure Specialists, and Product Experts across international teams to mitigate incidents and drive long-term solutions.
  • Build a deep understanding of Google's cloud technologies and distributed systems through an intensive training program covering technologies such as Borg, Colossus, Spanner, and other core GCP components.
  • Drive operational excellence by creating and maintaining technical documentation, standardizing incident response procedures, and continuously improving operational playbooks.
  • Lead and contribute to post-incident reviews, root cause analyses, and the implementation of preventive measures to improve platform reliability.
  • Identify opportunities for automation and contribute to improving operational efficiency, scalability, compliance, and service reliability.
  • Support the operation of highly secure cloud environments designed to meet stringent regulatory and sovereignty requirements.
We are looking forward to:
  • Several years of experience in Site Reliability Engineering, Cloud Operations, DevOps, Platform Engineering, Infrastructure Engineering, Production Support, Network Operations (NOC), Technical Operations, or a comparable role.
  • Experience operating and supporting business-critical production systems with demanding uptime and availability requirements.
  • Strong troubleshooting and incident management skills in complex technical environments.
  • Experience monitoring, operating, and maintaining distributed systems, cloud platforms, infrastructure services, or large-scale applications.
  • Familiarity with reliability engineering concepts, observability, monitoring, alerting, incident response, and root cause analysis.
  • Experience working with automation, scripting, operational tooling, or Infrastructure-as-Code approaches.
  • Strong analytical and problem-solving skills with a structured and methodical approach.
  • Professional proficiency in both German and English.
  • Willingness to participate in a regular on-call rotation.
  • Curiosity, adaptability, and a strong desire to learn and work with hyperscale cloud technologies.

The Group invests more than €4,5 billion per year in Research & Development in key areas, particularly for critical environments, such as Artificial Intelligence, cybersecurity, quantum and cloud technologies.

In 2025, the Group generated sales of €22.1 billion.

For our more than 85,000 employees in 65 countries we open up visionary perspectives, realise individual career paths and enable creative freedom. This is achieved with courage, versatility and the firm intention to make the demanding challenges of our time safer and more inclusive. With our sustainable value-focused management we support diversity actively.

*Human Intelligence

#LI-AF1

#LI-HYBRID

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

DevSecOps Engineer (m/w/d)
DevSecOps Engineer (m/w/d)

Thales • Deutschland

Hybrid
EUR 70.000 - 90.000
Field Chief Technology Officer (m/f/d) - Trusted Cloud Germany
Field Chief Technology Officer (m/f/d) - Trusted Cloud Germany

Thales • Berlin

Vor Ort
EUR 100.000 - 140.000
Site Reliability Manager GCP (w/m/d)
Site Reliability Manager GCP (w/m/d)

Thales Group • Berlin

Vor Ort
Field Chief Technology Officer (m/f/d) - Trusted Cloud Germany
Field Chief Technology Officer (m/f/d) - Trusted Cloud Germany

Thales Group • Berlin

Vor Ort
EUR 100.000 - 140.000
Creative freedom
Diversity-focused management
Visionary career paths
Senior Platform Engineer (m/f/d)
Senior Platform Engineer (m/f/d)

Thales Group • Berlin

Hybrid
EUR 70.000 - 90.000
Sovereign Cloud System Engineer (m/w/d)
Sovereign Cloud System Engineer (m/w/d)

Thales Group • Berlin

Hybrid
EUR 65.000 - 85.000
IT Infrastructure Operations Engineer (m/w/d)
IT Infrastructure Operations Engineer (m/w/d)

Thales Group • Berlin

Vor Ort
EUR 55.000 - 80.000
DevOps Engineer
DevOps Engineer

Thales Group • Berlin

Hybrid
EUR 60.000 - 85.000
DevSecOps Engineer (m/w/d)
DevSecOps Engineer (m/w/d)

Thales Group • Berlin

Hybrid
EUR 60.000 - 80.000
Technical Support Engineer (w/m/d)
Technical Support Engineer (w/m/d)

Thales • Berlin

Vor Ort
EUR 45.000 - 65.000