Senior Site Reliability Engineer

Real Work From Anywhere

Deutschland

Vor Ort

EUR 70.000 - 90.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Remote work
Generous PTO
Wellness and learning allowances
Annual Airalo Away retreat

Zusammenfassung

Tamarind Intelligence in Germany is seeking a Site Reliability Engineer to lead the design and implementation of scalable systems. You'll tackle complex challenges in a fully remote setup while ensuring high reliability across our global operations. Key responsibilities include leading post-incident reviews and working closely with software engineers to ensure systems are built for reliability.

The ideal candidate will have strong experience with AWS and Kubernetes, and good communication skills in English. We offer generous perks like remote work, wellness allowances, and a supportive workplace culture.

Qualifikationen

  • 5 years of experience as a Site Reliability Engineer or in a similar role.
  • 3 years of experience with AWS services including knowledge of container orchestration.
  • Experience with chaos engineering techniques for testing system resilience.

Aufgaben

  • Lead the design of scalable systems in a multi-region AWS environment.
  • Conduct blameless post-incident reviews to uncover root causes.
  • Work with software engineers to design systems for reliability and scalability.

Kenntnisse

Site Reliability Engineering
AWS Services
Kubernetes
Python
Observability Principles
Infrastructure as Code
CI/CD Tools

Ausbildung

Bachelor's degree in Computer Engineering or similar

Tools

Prometheus
Datadog
Terraform
GitHub Actions

Jobbeschreibung

Ready to make travel easier for millions? Airalo is the world's first and largest eSIM store, helping travellers stay connected seamlessly in over 200 countries and regions. We trust our teams to take ownership, put customers first, and do work that has a real impact every day.

What's in it for you?

Airalo offers team members a range of perks, including remote work, generous PTO, wellness and learning allowances, and, of course, our annual Airalo Away retreat.

About the Role

Hi, I'm Daniele, VP of Engineering at Airalo! Engineering drives Airalo's eSIM platform. We build the product that lets millions of people connect instantly across the globe. The challenges are exciting: high‑scale systems, carrier integrations, and products spanning both B2C and B2B. What matters most to us is creating an environment where engineers can do their best work. Real ownership, autonomy, and a direct link between what you ship and business outcomes are key. You'll work with smart, motivated people who take their craft seriously. If you want to build things that matter at global scale, this is where you do it.

Airalo's fully remote Engineering team in Spain is growing. In this role, you'll tackle complex technical challenges across our product ecosystem, helping build, innovate, and scale the platform that keeps millions of travellers connected worldwide.

We are a company that values SRE principles and practices. We empower our SREs to make data‑driven decisions, automate operational tasks, and continuously improve the reliability of our systems. We foster a blameless culture where everyone is encouraged to learn from mistakes and share knowledge. If you are passionate about building and maintaining highly reliable systems, we would love to hear from you!

On Call

Participating in our on‑call rotation is a core expectation of this role. It's essential for maintaining 24/7 service reliability across our global operations, ensuring our systems remain resilient and our customers experience uninterrupted service, regardless of time zone or geography.

  • Paid Rotation: We offer standby fees and overtime pay.
  • Delayed Start: No on‑call duties for your first six months.
  • Rest & Recovery: Guaranteed rest periods and flexible hours following night incidents.
  • Shared Load: Rotations are split (Weekdays vs. Weekends) to minimize fatigue.

For full details, refer to the On‑Call Policy in the Airalo Handbook.

What you'll do:
  • Lead the design of scalable, fault‑tolerant and self‑healing systems in a multi‑region AWS environment.
  • Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to drive architectural decisions and error budget policies.
  • Conduct blameless post‑incident reviews to uncover systemic root causes and implement long‑term preventive measures.
  • Identify patterns of manual work and lead the development of internal tools/automation to permanently eliminate them.
  • Develop and maintain automated runbooks and playbooks for common operational tasks and complex incident response.
  • Shift from simple monitoring to deep observability, ensuring high cardinality data leads to proactive actionable insights.
  • Proactively identify and mitigate operational risks through chaos engineering and architecture reviews.
  • Work with software engineers to design systems for reliability, scalability, and maintainability from the early stages of the SDLC.
  • Continuously evaluate and optimize system performance, capacity, and cost efficiency.
  • Beyond just participating, you will refine the on‑call experience to reduce alert fatigue, improve MTTR, and ensure sustainable rotation health.
Must‑haves:
  • Bachelor's degree in Computer Engineering or a similar discipline.
  • 5 years of experience as a Site Reliability Engineer or in a similar role.
  • 3 years of experience with AWS services including strong knowledge of container orchestration.
  • 2 years of Kubernetes experience.
  • Deep understanding of observability principles and tools such as Prometheus, Datadog, OpenTelemetry and similar.
  • Experience with leading incident management and complex postmortem analysis.
  • Experience and interest in managing infrastructure as code (Terraform).
  • Experience with chaos engineering and other techniques for testing system resilience.
  • Experience with CI/CD tools such as GitHub Actions for automated delivery.
  • Proficiency in at least one programming language (Python, Go, Java, etc.) for building automation and internal tooling.
  • Event‑driven architecture experience (SNS, SQS, etc.).
  • Ability to work independently and collaboratively in a fast‑paced environment.
  • Team player and open to new ideas.
  • Good communication skills and fluency in English.
Good to have:
  • Prior experience with Scrum and other agile methods.
  • Certification in relevant areas such as AWS Certified DevOps Engineer, Certified Kubernetes Administrator (CKA), or similar.
  • Prior experience with Telco Core Networks (5G/LTE Packet Core, IMS, Signaling) and low‑latency networking.
  • Experience with AI‑driven SRE tools for anomaly detection and improvements.
  • Contributions to open‑source SRE projects or communities.
  • Prior work experience in telecommunications.
  • Deep understanding of eSIM and GSMA related technologies and services.
Diversity & Inclusion

Airalo is an equal‑opportunity employer and values diversity, equity & inclusion. We do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We are committed to providing reasonable accommodations upon request for individuals with disabilities throughout our job interview process.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

iOS Developer
iOS Developer

Airalo • Deutschland

Remote
EUR 46.000 - 66.000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

1GLOBAL • Berlin

Vor Ort
EUR 90.000 - 140.000
Growth opportunities
International experience
Dynamic work environment
+2
Police Supervisor - Law Enforcement Expert
Police Supervisor - Law Enforcement Expert

Mercor • Berlin

Vor Ort
EUR 90.000 - 130.000
Senior Site Reliability Engineer (m/f/d) at TOPdesk
Senior Site Reliability Engineer (m/f/d) at TOPdesk

TOPdesk • Kaiserslautern

Vor Ort
EUR 90.000 - 120.000
Possibility to work remote
Flexible working hours
Senior Site Reliability Engineer (m/f/d)
Senior Site Reliability Engineer (m/f/d)

TOPdesk • Kaiserslautern

Vor Ort
EUR 90.000 - 150.000
30 days annual vacation
Remote-friendly options
Health and wellness programs
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Visa Hunt • Deutschland

Vor Ort
EUR 90.000 - 150.000
Home office budget
Learning & development budget of €1000
Competitive salary
+5
Site Reliability Engineer
Site Reliability Engineer

United States Digital Space LLC • Berlin

Hybrid
EUR 65.000 - 85.000
Senior Site Reliability Engineer (SRE / Backend) f/m/d
Senior Site Reliability Engineer (SRE / Backend) f/m/d

nilo • Berlin

Hybrid
EUR 110.000 - 150.000
Real ownership in a small team
Mental health platform access for you/
Free nilo app access (incl. family)
+7
Software Architect
Software Architect

Air Apps • Berlin

Vor Ort
EUR 80.000 - 100.000
Apple hardware
Annual Bonus
Top-tier Health and Life Insurance
+7
Senior Site Reliability Engineer (all genders)
Senior Site Reliability Engineer (all genders)

Meyandy LLC • Berlin

Hybrid
EUR 90.000 - 130.000
Hybrid work model