Senior Site Reliability Engineer (SRE)

spotme

Lausanne

Hybrid

CHF 150.000 - 190.000

Vollzeit

Vor 8 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

SpotMe is seeking a hands-on infrastructure expert to own the reliability of our 24/7 SaaS platform powering live events. You will design scalable cloud infrastructure, implement IaC with Terraform, and integrate AWS services to handle peak traffic and resilience.

You will lead on-call rotations, strengthen observability with Datadog and Jenkins, and automate deployments and patches. You’ll work with engineers to boost uptime, security, and cost efficiency.

Qualifikationen

  • Hands-on experience building scalable cloud infrastructure.
  • Strong incident response and on-call experience.
  • Proficiency with automation and infrastructure-as-code.

Aufgaben

  • Develop and deploy scalable infrastructure using Terraform and AWS.
  • Improve platform reliability and on-call incident response.
  • Automate provisioning and CI/CD pipelines; monitor and optimize performance.

Kenntnisse

Cloud infrastructure
Incident response
SRE principles
On-call experience
Automation

Tools

Terraform
AWS
Packer
Docker
Jenkins
Datadog
Locust
Gatling
Kubernetes

Jobbeschreibung

Mission - Why We Exist, What We Do, and Why We Need You

SpotMe is a leading B2B event platform that helps enterprises increase the impact of their events by delivering CRM-connected, high-quality experiences across in-person, virtual, hybrid events, and webinars. With a strong focus on life sciences, SpotMe powers Onomi, an HCP engagement product that enables medical and commercial teams to run impactful congresses, symposia, advisory boards, and webinars. Together, SpotMe and Onomi turn events into a company's most effective engagement channel. This role is for a hands-on technical expert who keeps high-performance, scalable systems running and who makes sure a 24/7 SaaS platform operates smoothly under all conditions. If you have a strong background in cloud infrastructure, automation, and incident response, this is your opportunity to take full ownership of our platform's reliability and scalability, in an environment where both are critical to the business. We build with AI-assisted development tools as a core part of how we work, and you'll have real latitude to use them. You will not just monitor the platform; you will drive lasting improvements to its uptime, performance, and resilience.

You will report to the Infrastructure Lead and work closely with the engineering and product teams. You will maintain and optimize the platform's infrastructure and build solutions that improve its reliability and scalability, making sure the platform scales smoothly to handle peak traffic and stays resilient during high-stakes live events. Your time will be spent on:

Your time will be spent on:
  • 40% Infrastructure development
    • Develop and deploy scalable infrastructure using Terraform and cloud-native AWS services.
    • Contribute to critical full-stack work that spans the back end and infrastructure or cloud development.
    • Build automation for infrastructure provisioning and CI/CD pipelines (Jenkins), including image builds with Packer and Docker.
  • 40% Platform optimization and resilience
    • Optimize the platform's cloud infrastructure for high availability and cost efficiency.
    • Monitor and update infrastructure to follow security best practices, applying necessary patches and upgrades.
    • Make sure the infrastructure can handle peak loads, scaling smoothly during high-traffic events.
  • 20% Support and observability
    • Take your turn in the on-call infrastructure rotation, responding to incidents and resolving them quickly.
    • Strengthen the platform's monitoring and observability (Datadog, Pingdom, Elastic) to catch issues before they reach end users.
    • Handle infrastructure support requests and drive continuous improvement in incident resolution.
Objectives - The Problems You Will Solve
In Your First Month:
  • Understand the current platform architecture and complete 3 infrastructure-as-code change request reviews.
  • Learn our monthly patching procedures and deploy critical security patches.
  • Get hands-on with our load testing framework (Locust, Gatling) and run one release-validation load test.
  • Take part in the weekly risk-analysis meeting and run a scheduled database scaling exercise.
  • Get set up with our AI-assisted development toolchain, including Anthropic's Claude, and use it in your daily work.
  • Handle and resolve at least 3 infrastructure support requests.
  • Build a report on what surprised you: what looked fragile, and what was hard to find documented.
After 3 Months:
  • Own a security-hardening improvement, such as tightening firewall and network rules across an environment.
  • Design and deliver one infrastructure-as-code project in Terraform, from proposal to production.
  • Build and ship one Python-based AWS Lambda that automates an operational task.
  • Reach the level of system knowledge needed to operate autonomously in the on-call rotation, resolving incidents without escalation.
After 6 Months:
  • Lead the resolution of a critical infrastructure incident and drive lasting improvements in incident response and recovery times.
  • Lead a significant reliability or tech-debt project, such as moving a complex on-premises build system to the cloud.
  • Identify repetitive engineering workflows and automate them end-to-end with AI-assisted tooling, so the team's time goes to the hard problems instead of the recurring ones.
  • Implement and own observability for one critical service end-to-end: instrumentation, alerting, and dashboards that let the team detect and diagnose issues before customers report them.
  • Own and measurably improve one reliability metric for a critical service, such as time to detect or mean time to recovery, against a baseline you establish in your first month, with the target agreed with your manager.
After 12 Months:
  • You have set the standard for how infrastructure is built here: your Terraform patterns, review practices, and documentation are what other engineers work from by default.
  • You own peak-traffic readiness end to end, with load testing, capacity review, and scaling runbooks running to a schedule you set and refine.

AI-ass

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer (SRE) New Lausanne, Switzerland (Hybrid)
Senior Site Reliability Engineer (SRE) New Lausanne, Switzerland (Hybrid)

SpotMe • Lausanne

Hybrid
CHF 140.000 - 210.000
Senior Site Reliability Engineer — Cloud, AI Tools, 24/7 Uptime
Senior Site Reliability Engineer — Cloud, AI Tools, 24/7 Uptime

spotme • Lausanne

Hybrid
CHF 150.000 - 190.000
Senior Cloud Reliability Engineer — AWS, Terraform, AI Tools
Senior Cloud Reliability Engineer — AWS, Terraform, AI Tools

SpotMe • Lausanne

Hybrid
CHF 140.000 - 210.000
Quant DevOps Engineer
Quant DevOps Engineer

SCOR • Zürich

Vor Ort
CHF 140.000 - 210.000
Site Reliability Engineer
Site Reliability Engineer

Embodied AI • Lausanne

Vor Ort
CHF 120.000 - 170.000
0D Capital: Senior Devops/SRE Engineer
0D Capital: Senior Devops/SRE Engineer

The10minutecareersolution • Genf

Remote
CHF 110.000 - 145.000
Competitive compensation
Equity/token upside
Remote work flexibility
Senior Technical Support Engineer, Observe by Snowflake
Senior Technical Support Engineer, Observe by Snowflake

Snowflake • Zürich

Vor Ort
CHF 90.000 - 110.000
Cloud Platform Lead 80 - 100 %
Cloud Platform Lead 80 - 100 %

Hamilton Medical AG • Ems

Hybrid
CHF 170.000 - 230.000
Quant DevOps Engineer
Quant DevOps Engineer

SCOR UK Company Limited • Zürich

Vor Ort
CHF 100.000 - 150.000
Applied AI Engineer
Applied AI Engineer

Cyber Resilience Shield • Zürich

Hybrid
CHF 110.000 - 140.000
CHF 4'000 yearly for work-related equipment
Team events including snowboarding and go-karting
Flexible work environment