Senior Site Reliability Engineer (80–100%)

Open Systems AG

Zürich

Vor Ort

CHF 140.000 - 190.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Open Systems AG is seeking a Senior Site Reliability Engineer to strengthen our Mission Control platform. You will design and implement automation, self service tooling, and reliable operations for production services running globally.

We expect senior software engineering skills, strong knowledge of distributed systems, and hands‑on experience with Kubernetes, Terraform, and observability stacks. Join a team focused on reliability, automation, and AI‑assisted tooling.

Qualifikationen

  • Strong software engineering background with production‑grade Go experience.
  • Solid understanding of distributed systems and scalable architecture.
  • Experience designing, building, and operating services and their APIs (REST/gRPC) in production.
  • Experience operating production systems, including incident response, on‑call, and root cause analysis.
  • Experience with SRE, DevOps, or platform engineering roles.
  • Hands‑on with infrastructure and operations tooling (Kubernetes, Terraform, GitOps, CI/CD).
  • Knowledge of Linux system administration, networking concepts, and common protocols.

Aufgaben

  • Build, evolve, and maintain automation frameworks for the Mission Control platform.
  • Develop self‑service APIs and tooling for customers and teams to operate services safely in production.
  • Define and improve SLIs, SLOs, error budgets, and SLA measurements across MC and services.
  • Own reliability initiatives from concept to production and long‑term operation.
  • Collaborate with AI Engineering to integrate AI‑assisted capabilities into automation.
  • Participate in incident response and lead root cause analyses to drive preventive actions.

Kenntnisse

Go (Golang)
Distributed systems
REST/gRPC APIs
Incident response
SRE/DevOps
Linux administration
Networking concepts
Communication skills
Proactive mindset

Ausbildung

University degree in Computer Science

Tools

Kubernetes
Terraform / IaC
GitOps
CI/CD tooling
Prometheus / Loki / Tempo
Cloud platforms (AWS/Azure/GCP)

Jobbeschreibung

Senior Site Reliability Engineer (80–100%)
Your Mission

We are seeking a highly motivated and talented Senior Site Reliability Engineer to join our growing team. As an SRE, you will be a key driver in building automation, reducing operational toil, and continuously improving how our services are operated, while ensuring reliability, scalability, and accurate SLA measurement.

You will combine strong software engineering and system architecture skills with an operations mindset to improve how our services are built and run at scale. Much of your work centers on the Mission Control (MC) platform— the operational backbone our engineers rely on to run customer services in production, and our point of interaction with customers, running 24/7 around the globe. To improve Mission Control, we build services such as an automation framework, a self‑service platform, and a service‑level monitoring system that eliminate toil and repetitive tasks, enable customers, and make operations more reliable and efficient. Your responsibilities will include:

  • Building Operational Automation: Design, build, and evolve the automation framework and tooling that powers the MC platform — primarily in Golang— with a strong focus on maintainability, scalability, and reliability. Beyond the framework itself, you will build automation workflows on top of the platform that reduce operational toil and improve day‑to‑day efficiency for the engineers and operators who run our services.
  • Developing Self‑Service APIs: Build and maintain the service APIs and self‑service operational tooling of the MC platform that enable customers and teams to safely and efficiently operate services in production without manual intervention.
  • Applying Site Reliability Engineering (SRE) Principles: Define, implement, and continuously improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and SLA measurements so that reliability is measurable and actionable across the MC platform and the services it operates.
  • Owning Reliability and Operations Initiatives: Take ownership of reliability, automation, and Mission Control projects, driving them independently from problem identification through implementation and long‑term operation.
  • Collaborating with AI Engineering: Work closely with our AI team and tooling, integrating AI‑assisted capabilities into our automation and operational workflows.
  • Incident Response and Learning: Participate in incident response and the on‑call rotation, leading root cause analysis and driving sustainable corrective and preventive actions.

This position will give you the opportunity to lead midsize to large automation and reliability initiatives, from early concept and design through production deployment and ongoing operations. You will collaborate closely with software engineers, platform teams, and product owners to identify and implement improvements across our services and operational practices. As part of the role, you will spend at least one day per week working within Mission Control, staying close to the day‑to‑day reality of how our services are run.

Your Qualifications

You are a senior engineer with a strong foundation in software development and system architecture, combined with a clear interest in operations and reliability. You are comfortable taking ownership, learning quickly, and driving change without needing detailed direction.

Required skills and qualifications:

  • Strong software engineering background, ideally with production‑grade Go (Golang) experience — proficiency in another modern language is fine if you are a quick learner and willing to adapt, as Go is our primary language.
  • Solid understanding of distributed systems and scalable architecture.
  • Proven experience designing, building, and operating services and their APIs (e.g. REST, gRPC) in production.
  • Experience operating production systems, including incident response, on‑call, and root cause analysis.
  • Experience with SRE, DevOps, platform engineering, or reliability‑focused roles.
  • Hands‑on experience with infrastructure and operations tooling, such as:
    • Kubernetes
    • Terraform / Infrastructure as Code
    • GitOps principles and CI/CD tooling
    • Prometheus, Loki, Tempo, and modern observability stacks
    • Major cloud platforms (AWS, Azure, GCP)
  • Knowledge of Linux system administration, networking concepts, and major Internet protocols (TCP/IP, IPsec, SSL, SSH, SMTP, HTTPS, DNS).
  • Ability to think in terms of systems, failure modes, and trade‑offs.
  • Strong communication skills and the ability to build trust across engineering and operations teams.
  • A proactive mindset: you see problems, propose solutions, and take responsibility for delivering them.
  • University degree in Computer Science or similar educational level.

Bonus skills and qualifications:

  • Experience building internal developer platforms or automations
  • Interest in or experience with integrating AI/LLM‑assisted capabilities into automation or operational tooling
  • Background in networking or security operations
  • Experience defining and operating SLO‑based reliability programs at scale
What We Offer

Want to join a team that enjoys making secure connectivity simple for our customers? You’ll be among people who believe in:

  • CaringPASSIONATELYabout keeping our customers safe – We’re dedicated to solving problems. Whatever it takes.
  • ThinkingUNCONVENTIONALLYto stay ahead – The world never fails to surprise us. So let’s surprise it first.
  • Doing the hard work to make thingsSIMPLE– Craft and hone something that delights in its simplicity.
  • WorkingCOLLABORATIVELYto build success – The power of the team will always make us faster and better.

As a testament to this, Open Systems has been recognized as an outstanding place to work. You’ll be surrounded by smart teams who enrich your experience and provide opportunities you will need to develop your skills and advance your career.

Come as you are! We are looking for amazing people of diverse backgrounds, experiences, abilities, and perspectives. Open Systems welcomes and encourages diversity in the workplace regardless of race, gender, religion, age, sexual orientation, disability, or veteran status.

About Open Systems

Open Systems is about secure connectivity made easy, combining SD-WAN, Firewall, SWG, CASB, and ZTNA into a framework that supports secure connectivity across cloud and hybrid environments and locations. Open Systems Managed SASE provides a comprehensive SASE solution through an easy‑to‑use customer portal, underpinned with a unified data platform to drive future innovation, all delivered as a 24x7 managed service.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Operations Engineer
Operations Engineer

Open Systems AG • Zürich

Hybrid
CHF 110.000 - 150.000
Passion for customer safety
Unconventional thinking
Simplicity in execution
+1
Senior Software Systems Engineer - Web Security
Senior Software Systems Engineer - Web Security

Open Systems • Zürich

Vor Ort
CHF 120.000 - 160.000
Operations Engineer
Operations Engineer

Open Systems • Zürich

Vor Ort
CHF 100.000 - 140.000
Customer safety
Unconventional thinking
Simplicity at work
+1
Senior Software Systems Engineer - Web Security
Senior Software Systems Engineer - Web Security

Open Systems AG • Zürich

Vor Ort
CHF 130.000 - 190.000
Software Systems Engineer - Web Security
Software Systems Engineer - Web Security

Open Systems • Zürich

Vor Ort
CHF 110.000 - 170.000
Software Systems Engineer - Web Security
Software Systems Engineer - Web Security

Open Systems AG • Zürich

Vor Ort
CHF 120.000 - 180.000
Senior SRE: Automation & AI-Driven Reliability Leader
Senior SRE: Automation & AI-Driven Reliability Leader

Open Systems AG • Zürich

Vor Ort
CHF 140.000 - 190.000
Senior Site Reliability Engineer (SRE) New Lausanne, Switzerland (Hybrid)
Senior Site Reliability Engineer (SRE) New Lausanne, Switzerland (Hybrid)

SpotMe • Lausanne

Hybrid
CHF 140.000 - 210.000
(Senior) Systems Engineer – Remote Access
(Senior) Systems Engineer – Remote Access

Open Systems • Zürich

Vor Ort
CHF 80.000 - 110.000
Cutting-Edge Technology
Global Impact
Career Growth
+1
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

spotme • Lausanne

Hybrid
CHF 150.000 - 190.000