Senior Site Reliability Engineer (SRE)

1GLOBAL

Berlin

Vor Ort

EUR 100.000 - 150.000

Vollzeit

Vor 12 Tagen
Bewerbungsgenerator

Eine maßgeschneiderte Bewerbung für diese Stelle — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Growth opportunities
International experience
Professional development
Dynamic work environment
Open communication
Tangible impact

Zusammenfassung

1GLOBAL is seeking a Senior Site Reliability Engineer (SRE) to join its Technology Department in Berlin, Germany. You will be responsible for reliability across cloud and on-prem environments, owning SLIs/SLOs, and leading incident response with a focus on automation and observability.

You will collaborate with DevOps, Infrastructure, IP Network, and Security teams to ensure carrier-grade reliability, design self-healing systems, and drive capacity planning and performance benchmarking for a

Qualifikationen

  • Minimum 5 years in Site Reliability, Systems, or Infrastructure Engineering (2+ years in SRE role).
  • Strong Linux, distributed systems, and networking expertise.
  • Experience building and running high-availability production systems.
  • Hands-on with redundancy testing, disaster recovery, and HA validation.
  • Deep understanding of monitoring, observability, and incident management.
  • Experience with Prometheus, Grafana, Loki, OpenTelemetry or similar.
  • Proficiency in Python, Go, and Bash for automation.
  • Knowledge of Kubernetes, container orchestration, and service mesh.
  • Experience with AWS (EKS, EC2, VPC) and on-prem integration.
  • IaC with Terraform.
  • Networking fundamentals (routing, load balancing, BGP, DNS, VXLAN).
  • Strong analytical, problem-solving, and cross-functional collaboration.
  • Excellent communication skills across distributed teams.

Aufgaben

  • Act as a senior technical contributor mentoring peers.
  • Define and maintain SLIs and SLOs for core infrastructure.
  • Plan redundancy and resilience testing across layers.
  • Design automated recovery mechanisms and self-healing workflows.
  • Lead incident response, root-cause analysis, and blameless post-mortems.
  • Develop and enhance observability using Prometheus, Grafana, Loki and OpenTelemetry.
  • Collaborate with Infra and DevOps for safe deployments and rollback policies.
  • Conduct fault injection, load, and chaos testing.
  • Reduce operational toil via automation and reliability tooling.
  • Contribute to on-call practices and runbooks.
  • Perform capacity planning, performance benchmarking, resilience audits.
  • Ensure security, reliability and availability standards.
  • Create internal docs, playbooks, and guidelines.
  • Contribute to cloud cost optimization initiatives.

Kenntnisse

Linux systems
Distributed systems
Networking
High-availability
Redundancy testing
Disaster recovery
Monitoring
Observability
Incident management
Prometheus
Grafana
Loki
OpenTelemetry
Python
Go
Bash
Kubernetes
Service mesh
AWS
Terraform
Networking fundamentals
Analytical skills
Communication skills

Tools

Terraform
AWS (EKS, EC2, VPC)
On-prem integration

Jobbeschreibung

1GLOBAL is a technology-driven global mobile communications provider helping enterprises and consumer brands deliver seamless connectivity worldwide. Powered by a best-in-class telecom platform, proprietary eSIM technology, and its own mobile core network across 15 countries, 1GLOBAL operates as a fully regulated telecommunications provider in 43 countries.

Founded in 2022, 1GLOBAL has rapidly become one of Europe's fastest-growing telecom technology companies, connecting more than 80 million people and devices globally. Our customers include leading banks, multinational enterprises, global retailers, travel companies, payment service providers, and digital-first businesses.

Headquartered in the Netherlands, with R&D hubs in Lisbon, Berlin, and São Paulo, 1GLOBAL employs over 550 professionals across 16 countries. With annual revenues exceeding US$200 million and a strong track record of profitability, we continue to invest in innovation, infrastructure, and international expansion as we redefine the future of global mobile connectivity.

About the Ideal Candidate

We are looking for a talented Senior Site Reliability Engineer (SRE) to join our Technology Department. We are open to hiring this role in Berlin, Germany.

As a Senior SRE, you will be a senior individual contributor responsible for strengthening the stability, scalability, and reliability of our global infrastructure and services across both cloud and on-prem environments. You will work alongside SREs under the guidance of the SRE Team Lead, taking ownership of critical reliability domains and helping drive a data-driven reliability culture based on SLIs, SLOs, and error budgets.

Your mission will be to proactively identify weaknesses across systems and improve reliability through redundancy testing, automation, and observability. You will design, build, and operate the tools and processes that automatically detect, prevent, and recover from incidents, ensuring our services remain reliable and performant for customers around the world.

This role collaborates closely with DevOps, Infrastructure, IP Network, and Security teams to maintain carrier-grade reliability standards across all layers of our platform.

About the Role
  • Act as a senior technical contributor within the SRE team, mentoring peers and setting the technical bar for reliability engineering.
  • Define, measure, and maintainSLIs and SLOs for core infrastructure and customer-facing services.
  • Plan and executeredundancy and resilience testing across service, infrastructure, and networking layers — validating failover, HA configurations, and disaster recovery readiness.
  • Design and implementautomated recovery mechanisms, self-healing workflows, and intelligent alerting systems.
  • Driveincident response, root-cause analysis, and blameless post-mortems, and ensure implementation and tracking of corrective and preventive actions derived from them to achieve continuous improvement.
  • Develop and enhance observability (metrics, logs, traces) using Prometheus, Grafana, Loki, and OpenTelemetry.
  • Partner with Infrastructure and DevOps teams to ensure deployment safety, rollback policies, and configuration consistency.
  • Proactively identify weaknesses through fault-injection, load, and chaos testing.
  • Continuously reduce operational toil through automation and reliability tooling.
  • Contribute to on-call practices, improving alert quality, runbooks, escalation procedures, and incident management processes.
  • Perform capacity planning, performance benchmarking, and resilience audits across systems.
  • Ensure compliance with security, reliability, and availability standards.
  • Create and maintain internal documentation, playbooks, and operational guidelines for peers and users.
  • Contribute to cloud cost-optimization initiatives, including reserved capacity planning, autoscaling design, storage tiering, workload right-sizing, and continuous anomaly detection.
About Us

1GLOBAL is a technology-driven global mobile communications provider helping enterprises and consumer brands deliver seamless connectivity worldwide. Powered by a best-in-class telecom platform, proprietary eSIM technology, and its own mobile core network across 15 countries, 1GLOBAL operates as a fully regulated telecommunications provider in 43 countries.

Founded in 2022, 1GLOBAL has rapidly become one of Europe's fastest-growing telecom technology companies, connecting more than 80 million people and devices globally. Our customers include leading banks, multinational enterprises, global retailers, travel companies, payment service providers, and digital-first businesses.

Headquartered in the Netherlands, with R&D hubs in Lisbon, Berlin, and São Paulo, 1GLOBAL employs over 550 professionals across 16 countries. With annual revenues exceeding US$200 million and a strong track record of profitability, we continue to invest in innovation, infrastructure, and international expansion as we redefine the future of global mobile connectivity.

Requirements
About You
  • A minimum of 5 years of experience in Site Reliability, Systems, or Infrastructure Engineering (including 2+ years in a dedicated SRE role).
  • Strong expertise in Linux systems engineering, distributed systems, and networking.
  • Proven experience building and running high-availability, mission-critical production systems.
  • Hands-on experience with redundancy and failover testing, disaster recovery, and high-availability architecture validation.
  • Deep understanding of monitoring, observability, and incident management principles
  • Experience with Prometheus, Grafana, Loki, Thanos, and OpenTelemetry or similar tools.
  • Proficiency in Python, Go, and Bash for automation and reliability tooling.
  • Strong knowledge of Kubernetes, container orchestration, and service mesh architectures.
  • Experience with AWS (EKS, EC2, VPC) and on-premises infrastructure integration.
  • Proficiency in Infrastructure as Code tools such as Terraform.
  • Understanding of networking fundamentals (routing, load balancing, BGP, DNS, VXLAN, etc.).
  • Excellent analytical and problem-solving skills, capable of operating under pressure.
  • Strong communication and collaboration skills across distributed and cross-functional teams
Nice-to-haves
  • Experience in telecom, carrier-grade, or large-scale distributed systems environments
  • Hands-on experience with chaos engineering and automated failure-scenario validation (e.g., simulating link or node failures)
  • Strong understanding of high-availability networking concepts
  • Background in capacity planning, traffic engineering, and multi-region failover
  • Experience building reliability dashboards and integrating SRE metrics into business KPIs or compliance reports
  • Familiarity with security and resilience standards (ISO 27001, NIST SP 800-53)
Benefits
Why 1GLOBAL?
  • Growth Opportunities: Advance your career in one of the fastest growing telecommunications companies, expanding over 50% year-on-year under the leadership of successful tech entrepreneurs
  • Major Transaction Exposure: Be in the driver's seat for transactions that will have an impact on the future telco industry
  • Work with a Talented Team: From the Board and the Founders to the Senior Management Team, you will collaborate daily with the most capable and renowned external advisors and constantly being exposed to talented and driven individuals
  • Dynamic Work Environment: Thrive in a collaborative, fast-paced workplace where innovation is encouraged, and every contribution counts
  • Professional Development: Work alongside industry experts to enhance your skills and knowledge in a cutting-edge field
  • International Experience: Gain opportunities to work in different 1GLOBAL offices around the world as you grow within the company
  • Open Communication Culture: Join a team where your ideas are heard, and open dialogue is encouraged, fostering a supportive and transparent work environment
  • Get Things Done Attitude: Be part of a results-driven team that values efficiency, creativity, and the drive to make a tangible impact in the industry

1GLOBAL is an equal opportunity employer, we value your character as much as your talent. Diversity drives our innovation, and we offer a collaborative, dynamic, and international work environment. We are excited for you to join our mission to revolutionise connectivity globally.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior DevSecOps Engineer - Berlin / Lisbon / Warsaw
Senior DevSecOps Engineer - Berlin / Lisbon / Warsaw

1GLOBAL • Berlin

Vor Ort
EUR 60.000 - 85.000
Growth opportunities
International experience
Professional development
+1
Senior Java Software Engineer
Senior Java Software Engineer

1GLOBAL • Berlin

Vor Ort
EUR 70.000 - 110.000
Growth opportunities
International experience
Open communication culture
+3
Data Engineer
Data Engineer

1GLOBAL • Berlin

Vor Ort
EUR 70.000 - 110.000
Growth opportunities
Major transaction exposure
Work with a talented team
+5
Sales Support Specialist (German-speaking)
Sales Support Specialist (German-speaking)

1global • Berlin

Vor Ort
EUR 42.000 - 64.000
Java Software Engineer
Java Software Engineer

1GLOBAL • Berlin

Vor Ort
EUR 50.000 - 70.000
Growth Opportunities
Major Transaction Exposure
Work with a Talented Team
+5
Senior Java Software Engineer
Senior Java Software Engineer

Meyandy LLC • Berlin

Vor Ort
EUR 70.000 - 110.000
Growth opportunities
Engaging team culture
Senior AI Engineer
Senior AI Engineer

1global • Berlin

Hybrid
EUR 120.000 - 180.000
Business & Data Analyst - B2B2C
Business & Data Analyst - B2B2C

1GLOBAL • Berlin

Vor Ort
EUR 65.000 - 90.000
Growth opportunities
International experience
Professional development
+1
Senior AI Product Manager
Senior AI Product Manager

1global • Berlin

Hybrid
EUR 90.000 - 130.000
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FACT-Finder • Berlin

Vor Ort
EUR 110.000 - 150.000
Hybrid work model
Flexible work policy
AI-driven environment