Streaming Reliability Engineer

XpertDirect

Stockholms kommun

On-site

SEK 900,000 - 1,200,000

Full time

25 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Mobility Technology in Stockholm seeks a Streaming Reliability Engineer to own the performance, resilience, and operability of high-throughput streaming infrastructure processing real-time data.

You will work at the intersection of Streaming Engineering, SRE, and Platform Engineering to keep Kafka and Flink workloads fast, observable, and reliable as data volumes grow.

Qualifications

  • 4+ years in Streaming Engineering, SRE, Data Infrastructure or Platform Engineering.
  • Proficient in Java and/or Scala.
  • Strong understanding of distributed systems and production reliability.

Responsibilities

  • Operate and improve high-throughput Kafka infrastructure.
  • Build and optimise real-time processing workloads using Apache Flink.
  • Improve reliability across Kubernetes-based streaming environments.
  • Define SLIs, SLOs, and reliability standards for critical streaming services.
  • Build monitoring and alerting using Prometheus.
  • Investigate latency, throughput, consumer lag, backpressure, and processing failures.
  • Automate infrastructure provisioning and configuration using Terraform.
  • Improve partitioning, scaling, and resource-allocation strategies.
  • Build reliability tooling in Java and/or Scala.
  • Automate recovery and reduce manual intervention during incidents.
  • Perform capacity planning for growing event volumes.
  • Partner with Data and Platform Engineers to design more resilient streaming systems.

Skills

Streaming engineering experience
Java/Scala
Distributed systems

Tools

Terraform
Prometheus
Kubernetes
Kafka
Flink

Job description

Mobility Technology | Streaming Infrastructure | Site Reliability Engineering | Real-Time Data | Distributed Systems

Our client, a growing Mobility Technology company based in Stockholm, is looking for a Streaming Reliability Engineer to own the performance, resilience, and operability of high-throughput streaming infrastructure processing real-time operational data.

You'll work at the intersection of Streaming Engineering, SRE, and Platform Engineering, ensuring Kafka and Flink workloads remain fast, observable, and reliable as data volumes and platform complexity grow.

What You'll Work On
  • Operate and improve high-throughput Kafka infrastructure
  • Build and optimise real-time processing workloads using Apache Flink
  • Improve reliability across Kubernetes-based streaming environments
  • Define SLIs, SLOs, and reliability standards for critical streaming services
  • Build monitoring and alerting using Prometheus
  • Investigate latency, throughput, consumer lag, backpressure, and processing failures
  • Automate infrastructure provisioning and configuration using Terraform
  • Improve partitioning, scaling, and resource-allocation strategies
  • Build reliability tooling in Java and/or Scala
  • Automate recovery and reduce manual intervention during incidents
  • Perform capacity planning for growing event volumes
  • Partner with Data and Platform Engineers to design more resilient streaming systems
Core Skills
  • 4+ years in Streaming Engineering, SRE, Data Infrastructure, Platform Engineering, or similar roles
  • Java and/or Scala
  • Strong understanding of distributed systems and production reliability
Nice to Have
  • Kafka Streams / Kafka Connect
  • Exactly-once processing concepts
  • Event-driven architectures
  • AWS / GCP
  • JVM performance tuning
  • Incident management and postmortems
  • High-throughput or low-latency systems
  • Experience operating multi-region streaming infrastructure
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Kafka Support Engineer
Kafka Support Engineer

Bytespoke • Stockholms kommun

On-site
SEK 600,000 - 800,000
Site Reliability Engineer
Site Reliability Engineer

Linuxcareers • Stockholms kommun

On-site
SEK 600,000 - 800,000
Kafka Reliability Engineer
Kafka Reliability Engineer

Bytespoke • Stockholms kommun

On-site
SEK 600,000 - 800,000
Senior Software Engineer – Backend Platform
Senior Software Engineer – Backend Platform

Evolution Gaming Limited • Stockholms kommun

On-site
SEK 550,000 - 700,000
Senior Kafka Platform Engineer
Senior Kafka Platform Engineer

emagine • Stockholms kommun

Hybrid
SEK 900,000 - 1,300,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Nordic Investin Group Aktiebolag • Stockholms kommun

On-site
SEK 900,000 - 1,100,000
Site Reliability Engineer
Site Reliability Engineer

Segment (Twilio) • Göteborgs kommun

On-site
SEK 900,000 - 1,200,000
Scala/Java Developer
Scala/Java Developer

Functional Software Stockholm AB • Malmö kommun

On-site
SEK 600,000 - 900,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Swediumglobal • Stockholms kommun

Hybrid
Platform SRE: Scalable Infra, CI/CD & Observability
Platform SRE: Scalable Infra, CI/CD & Observability

Linuxcareers • Stockholms kommun

On-site
SEK 600,000 - 800,000