Site Reliability Engineer (m/f/d)

Ververica | Original creators of Apache Flink®

München

Vor Ort

EUR 60.000 - 80.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

A leading technology firm in Munich seeks a Site Reliability Engineer to maintain infrastructure across AWS, GCP, and Azure. The role involves collaborating with teams, optimizing performance, and implementing SRE principles. Candidates should have a Bachelor's in Computer Science and experience with Kubernetes and Terraform. This is a full-time position offering opportunities in a dynamic environment.

Qualifikationen

  • Minimum 2 years of hands-on experience with Kubernetes clusters.
  • Proficiency in designing and maintaining Terraform code.
  • Strong knowledge of observability tools and practices.

Aufgaben

  • Build and maintain the infrastructure for Unified Streaming Data Platform.
  • Design and manage Infrastructure as Code using Terraform.
  • Implement and enhance observability tooling.

Kenntnisse

Kubernetes
Terraform
Observability tools
Linux systems
Cloud-native security

Ausbildung

Bachelor's degree in Computer Science or related field

Jobbeschreibung

Ververica, founded by the original creators of Apache Flink, empowers businesses to unlock the full potential of real-time data processing and analytics. Our platform provides cutting-edge stream processing and event-driven applications, enabling companies worldwide to build scalable and reliable data-driven solutions.

Role Overview

As a Site Reliability Engineer (SRE) at Ververica, you will design, provision, and maintain the infrastructure for Ververica's Unified Streaming Data Platform across multiple cloud providers, including AWS, GCP, and Azure. You will collaborate with software engineering teams to develop solutions that enhance feature delivery, optimize performance, and address security vulnerabilities. Your role will involve architectural improvements, implementation ownership, and driving reliability best practices.

Key Responsibilities

  • Build and maintain the infrastructure for Ververica's Unified Streaming Data Platform across AWS, GCP, and Azure
  • Design and manage Infrastructure as Code (IaC) using Terraform, ensuring modularity, reusability, and best practices
  • Implement and enhance observability tooling, including Grafana, Prometheus, logging systems, traces, metrics, dashboards, and alerts
  • Ensure system reliability through SRE best practices, including defining SLIs, SLOs, and error budgets
  • Improve infrastructure architecture and engineering efficiency through continuous evaluation and optimization
  • Enhance CI/CD pipelines to automate development workflows
  • Monitor, identify, and resolve security vulnerabilities (CVE updates and security enhancements)
  • Contribute to the successful development and launch of new products, features, and services
  • Periodically participate in on-call rotations to manage incidents in a 24/7 live infrastructure
  • Maintain and update documentation, including architectural designs and changes

About Ververica

Ververica, founded by the original creators of Apache Flink, empowers businesses to unlock the full potential of real-time data processing and analytics. Our platform provides cutting-edge stream processing and event-driven applications, enabling companies worldwide to build scalable and reliable data-driven solutions.

Role Overview

As a Site Reliability Engineer (SRE) at Ververica, you will design, provision, and maintain the infrastructure for Ververica's Unified Streaming Data Platform across multiple cloud providers, including AWS, GCP, and Azure. You will collaborate with software engineering teams to develop solutions that enhance feature delivery, optimize performance, and address security vulnerabilities. Your role will involve architectural improvements, implementation ownership, and driving reliability best practices.

Key Responsibilities

  • Build and maintain the infrastructure for Ververica's Unified Streaming Data Platform across AWS, GCP, and Azure
  • Design and manage Infrastructure as Code (IaC) using Terraform, ensuring modularity, reusability, and best practices
  • Implement and enhance observability tooling, including Grafana, Prometheus, logging systems, traces, metrics, dashboards, and alerts
  • Ensure system reliability through SRE best practices, including defining SLIs, SLOs, and error budgets
  • Improve infrastructure architecture and engineering efficiency through continuous evaluation and optimization
  • Enhance CI/CD pipelines to automate development workflows
  • Monitor, identify, and resolve security vulnerabilities (CVE updates and security enhancements)
  • Contribute to the successful development and launch of new products, features, and services
  • Periodically participate in on-call rotations to manage incidents in a 24/7 live infrastructure
  • Maintain and update documentation, including architectural designs and changes


Requirements

  • Bachelor's degree in Computer Science, Information Technology, or a related field
  • Minimum 2 years of hands-on experience with Kubernetes clusters, Helm charts, controllers, and operators
  • Proficiency in designing and maintaining Terraform code with best practices
  • Strong knowledge of observability tools and practices, including metrics, logging, and alerting systems
  • Experience implementing SRE principles such as SLIs, SLOs, and error budgets
  • Solid understanding of Linux systems and networking in cloud environments
  • Hands-on experience managing multiple Kubernetes clusters
  • Familiarity with distributed systems or streaming data platforms
  • Knowledge of cloud-native security best practices
Seniority level
  • Seniority level
    Mid-Senior level
Employment type
  • Employment type
    Full-time
Job function
  • Job function
    Other
  • Industries
    IT Services and IT Consulting

Referrals increase your chances of interviewing at Ververica | Original creators of Apache Flink by 2x

Sign in to set job alerts for “Site Reliability Engineer” roles.
Senior Site Reliability Engineer (w/m/d)
Senior Site Reliability / Gitops Engineer
Software Engineer (Python/Linux/Packaging)
Python and Kubernetes Software Engineer - Data, AI/ML & Analytics
Frontend software engineer (React) - Europe Remote
Senior Software Development Engineer in Test
Software Engineer – Junior bis Senior (all levels) / Banking (all genders)
Software Engineer - Python - Container Images
Software Engineer - Python - Container Images
Software Engineer - Python - Container Images
Python and Kubernetes Software Engineer - Data, Workflows, AI/ML & Analytics
Python Backend Senior Software Engineer - Remote 4 days a week (Europe)

We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Site Reliability Engineer (all genders)
Site Reliability Engineer (all genders)

TieTalent • Dresden

Hybrid
EUR 65.000 - 85.000
Training & Development
Altersvorsorge
Shopping Rabatte
+2
Senior Solutions Architect (m/f/d)
Senior Solutions Architect (m/f/d)

Ververica | Original creators of Apache Flink® • Frankfurt

Vor Ort
EUR 80.000 - 100.000
Senior Software Engineer (Flink Ecosystem)
Senior Software Engineer (Flink Ecosystem)

Ververica | Original creators of Apache Flink® • Berlin

Vor Ort
EUR 60.000 - 80.000
Site Reliability Engineer
Site Reliability Engineer

Contorion • Berlin

Hybrid
EUR 60.000 - 80.000
Flexible schedule
30 days of vacation
Subsidized job ticket or bike
+2
Java / Kotlin Backend-Entwickler (w/m/d)
Java / Kotlin Backend-Entwickler (w/m/d)

TieTalent • Leipzig

Hybrid
EUR 60.000 - 80.000
Mobile Working
30 Tage Urlaub
Versicherungsleistungen
+1
Senior Elixir Software Engineer (all genders)
Senior Elixir Software Engineer (all genders)

Distribusion Technologies • Berlin

Remote
EUR 70.000 - 90.000
Flexible working conditions
Relocation opportunities
Career growth
(Senior) Kotlin / Java Software Engineer - Client Technology (m/f/x) (onsite / remote in Germany)
(Senior) Kotlin / Java Software Engineer - Client Technology (m/f/x) (onsite / remote in Germany)

Scalable Capital • München

Remote
EUR 70.000 - 80.000
Attractive compensation package
Monthly contribution for transportation
Flexible vacation policy
+4
Senior Fullstack Engineer – Backend Focus (m/f/d)
Senior Fullstack Engineer – Backend Focus (m/f/d)

remberg • München

Hybrid
EUR 65.000 - 90.000
Competitive compensation package
Health and wellness benefits
Work-from-anywhere flexibility
+2
Full-Stack-Developer (m/f/d)
Full-Stack-Developer (m/f/d)

JustRelate • Deutschland

Remote
EUR 85.000 - 146.000
Remote work options
AWS cloud trainings and certifications
Language courses during working hours
+3
Python and Kubernetes Software Engineer - Data, AI/ML & Analytics
Python and Kubernetes Software Engineer - Data, AI/ML & Analytics

Canonical • Frankfurt

Remote
EUR 70.000 - 90.000
Personal learning and development budget
Annual compensation review
Recognition rewards
+2