Cloud Performance Engineering - Site Reliability Engineer

Smile Digital Health

Deutschland

Vor Ort

EUR 90.000 - 125.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Remote Work Environment
Flexible Time Away From Work Policy
Competitive Salary and Health/Medical
RRSP/TFSA/401K Employee Contribution
Life and Disability
Employee Assistance Program
FHIR Study Program and Skillsoft
Super HAPI Fun Club

Zusammenfassung

Smile Digital Health is looking for a Cloud Site Reliability Engineer to ensure reliability, scalability, and performance of production services across Azure and other cloud platforms. You will design and automate performance testing, integrate them into CI/CD, and use observability tools to proactively detect bottlenecks, collaborating with engineering, product, and security to meet SLAs.

You will implement secure, scalable cloud infrastructure, develop cost-tracking, and mentor teams on

Qualifikationen

  • 5+ years of experience with Cloud Service Providers and best practices around implementation and configuration, preferably Azure environments.
  • Experience across multiple cloud providers (Azure required; AWS/GCP an asset).
  • Hands-on performance testing and observability across distributed cloud-native apps.

Aufgaben

  • Collaborate with Security Operations to define cloud configuration best practices for Azure and other providers.
  • Develop multi-tenant approaches for databases, container platforms, auth, certificates, and registries.
  • Design, develop, and maintain cloud performance testing strategies and environments for scalability and resiliency.

Kenntnisse

Cloud performance engineering
Chaos engineering
Observability
OpenTelemetry
Prometheus
Grafana

Ausbildung

Bachelor's degree in CS/Engineering

Tools

Docker
Kubernetes
Terraform
JMeter
Gatling
Azure Load Testing
k6
Kafka
Application Insights
Log Analytics

Jobbeschreibung

Working for a company like Smile Digital Health means supporting our mandate for #BetterGlobalHealth. We strive towards this goal every day, and the results can be seen in the impact of our innovative health data platform and data management solutions, which are used in over 20 countries. We were #19 on Deloitte's Technology Fast 50 Ranking for 2024!

Smile Digital Health makes it easy for healthcare stakeholders to collect and exchange data with our leading FHIR-based data liberation platform.

At its heart, the Smile platform enables people and organizations to better manage healthcare data. We help generate and liberate structured healthcare data to ensure effective delivery across care teams and health systems bringing #BetterGlobalHealth to patients everyday!

The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure platforms for Smile Digital Health, its clients, and partners.

This role designs and automates performance testing frameworks, integrates them into CI/CD pipelines, and uses observability tools to proactively detect and resolve bottlenecks. Working closely with engineering, product, and security teams, the SRE ensures systems meet strict SLAs for performance and availability while driving continuous optimization across multiple cloud platforms.

Responsibilities:
  • Collaborate with our Security Operations teams to define and implement best practices around Cloud Service Provider configuration for Azure and other cloud providers.
  • Develop, implement, and coordinate a multi-tenant approach around service offerings for databases, container platforms, authentication, certificates, and product registries.
  • Design, develop, and maintain cloud performance testing strategies, frameworks, and environments to validate application scalability, reliability, and resiliency.
  • Develop and automate load, stress, spike, and endurance (soak) testing as part of CI/CD pipelines.
  • Analyze application and infrastructure performance to identify bottlenecks and recommend performance optimizations across cloud-native services.
  • Develop and maintain cost and utilization tracking and attribution processes across Cloud Service Providers.
  • Create documentation detailing Cloud Service Provider offerings, implementation patterns, and best practices.
  • Develop and maintain technical relationships with our core Cloud Service Providers.
  • Implement and maintain secure, scalable infrastructure platforms for delivering cloud services.
  • Ensure internal and external SLAs are consistently met or exceeded, while continuously monitoring and improving system performance, reliability, and availability.
  • Create tools for automating deployment, monitoring, and platform operations.
  • Implement and manage observability solutions (logging, metrics, tracing) using OpenTelemetry, Prometheus, Grafana, Azure Monitor, and related technologies to provide actionable performance insights.
  • Plan and execute chaos engineering experiments to evaluate and improve application resiliency and fault tolerance.
Requirements:
  • 5+ years of experience with Cloud Service Providers and best practices around implementation and configuration, preferably managing Azure environments supporting SaaS products.
  • Experience working across multiple cloud providers (Azure required; AWS and/or Google Cloud Platform considered an asset).
  • Strong experience in Cloud Performance Engineering, including performance analysis, capacity planning, scalability testing, and optimization of distributed cloud-native applications.
  • Proven experience working with microservices architecture, with a strong focus on Java-based services.
  • Experience applying Chaos Engineering practices to evaluate and improve system resiliency.
  • Strong experience designing and executing performance testing strategies, including load, stress, spike, and endurance (soak) testing, to validate application scalability and defined latency and error-rate thresholds.
  • Hands-on experience with performance testing tools such as JMeter, Gatling, Azure Load Testing, or k6.
  • Experience validating application services sustaining 500+ transactions per second (TPS) while meeting defined performance objectives.
  • Hands-on experience deploying and managing containerized applications using Docker and Kubernetes, including autoscaling and performance optimization.
  • Experience using Terraform to provision and manage cloud infrastructure using Infrastructure as Code (IaC).
  • Experience tuning Kafka (partitioning, consumer group sizing, throughput/latency trade-offs) and other messaging/queueing platforms to sustain target transaction rates.
  • Hands-on experience implementing and using observability platforms including OpenTelemetry, Prometheus, Grafana, Azure Monitor, Application Insights, and Log Analytics.
  • Proven experience with Security and Compliance (SOC 2, HIPAA, ISO 27001) best practices and implementing controls that support high-velocity software delivery teams.
Benefits:
  • Remote Work Environment
  • Flexible Time Away From Work Policy including PTO, Personal and Sick Days
  • Competitive Salary and Health/Medical Benefits
  • RRSP/TFSA/401K Employee Contribution
  • Life and Disability
  • Employee Assistance Program
  • FHIR Study Program and Skillsoft Learning
  • Super HAPI Fun Club

Smile's core values include respect, inclusion, embracing our differences, and celebrating shared values because our people are the foundation of our success. We are big on creating a sense of belonging and empowering each other to bring our authentic selves to work. We are dedicated to fostering a workplace that values diversity, equity, and inclusion.

We welcome and encourage candidates of all backgrounds to apply. Candidates are encouraged to inform us if they wish to discuss or require accommodations during interviews or while working at Smile.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Director, Cloud & Managed Services
Director, Cloud & Managed Services

Jobtailor • Deutschland

Hybrid
EUR 140.000 - 220.000
Senior Cloud Engineer (Fully Remote)
Senior Cloud Engineer (Fully Remote)

Smile Digital Health • Deutschland

Remote
EUR 90.000 - 140.000
Flexible time off
Remote work
Competitive salary
+3
Senior Site Reliability Engineer (x/f/m)
Senior Site Reliability Engineer (x/f/m)

United States Digital Space LLC • Berlin

Hybrid
EUR 90.000 - 140.000
Deutschlandticket
Vacation days
Health insurance
+7
Engineering Manager - Site Reliability & Observability (x/f/m)
Engineering Manager - Site Reliability & Observability (x/f/m)

United States Digital Space LLC • Berlin

Hybrid
EUR 120.000 - 180.000
Deutschlandticket (Germany-wide public
health insurance
Pension scheme (bAV)
+4
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Deutschland

Hybrid
EUR 126.000 - 140.000
Competitive Pay
Bonus Program
Benefits Package
+3
Systems Engineer
Systems Engineer

Cloudflare • Deutschland

Hybrid
EUR 100.000 - 140.000
Equity plan
Senior Site Reliability Engineer (m/f/d)
Senior Site Reliability Engineer (m/f/d)

TOPdesk • Kaiserslautern

Vor Ort
EUR 90.000 - 150.000
30 days annual vacation
Remote-friendly options
Health and wellness programs
+2
Staff Site Reliability Engineer (x/f/m)
Staff Site Reliability Engineer (x/f/m)

Doctolib • Berlin

Hybrid
EUR 110.000 - 150.000
Deutschlandticket (Germany-wide travel
28 vacation days
Hybrid work policy
+6
Senior Site Reliability Engineer (x/f/m)
Senior Site Reliability Engineer (x/f/m)

Doctolib • Berlin

Hybrid
EUR 90.000 - 140.000
Health insurance
Vacation days
Remote work flexibility
+1
Senior Software Engineer, .NET - US Remote
Senior Software Engineer, .NET - US Remote

Perfectserve • Deutschland

Remote
EUR 115.000 - 140.000
Health, Dental, Vision Insurance
401K with company match
17 company holidays plus paid time off