Senior Software Engineer - Reliability, Infrastructure, and Tooling

Jobgether

Deutschland

Remote

EUR 117.000 - 261.000

Vollzeit

Vor 3 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Eine vollständige Bewerbung in einer Minute — maßgeschneiderter Lebenslauf und Anschreiben, fertig zum Versenden.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Fully remote work
Equity participation
Health, dental, and vision benefits
Flexible vacation policy
Open-source collaboration
Real-time AI infrastructure exposure
Shared on-call practices
Equal opportunity employment

Zusammenfassung

Jobgether is seeking a Senior Software Engineer — Reliability, Infrastructure, and Tooling to design and operate reliable, scalable infrastructure for real-time and AI workloads. You will work closely with product teams to build secure, observable, and resilient systems at global scale.

You will tackle complex challenges in distributed architectures, implement instrumentation, and evolve tooling to empower developers.

Qualifikationen

  • Strong experience building and operating production systems with high concurrency and distributed workloads.
  • Significant experience with Kubernetes or equivalent large-scale container orchestration.
  • Deep knowledge of Linux internals and networking to troubleshoot across layers.
  • Experience using observability, monitoring, and logging tools to diagnose production issues.
  • Experience operating large-scale, globally distributed systems and managing configuration and technical debt.
  • Experience handling complex production incidents and implementing durable fixes.
  • Experience with open-source infrastructure technologies such as Kafka, ClickHouse, or comparable distributed systems.
  • Strong systems-thinking, able to reason about infrastructure in terms of signals, feedback, dependencies, and control mechanisms.
  • Excellent communication and collaboration with partner engineering teams.
  • Pragmatic engineering balancing delivery with maintainability and cost.
  • Interest in observability, reliability engineering, clean configuration, automation, and reducing operational complexity.
  • Nice to have: data engineering and analytics.
  • Nice to have: global Layer 3 networking.
  • Nice to have: operating systems with long-lived workloads such as real-time media.
  • Nice to have: Google SRE or another high-scale reliability environment.
  • Nice to have: PCI compliance experience.

Aufgaben

  • Ramp up on a complex global architecture involving distributed databases, messaging systems, networking infrastructure, Kubernetes, and other core platform technologies, identifying areas of reliability debt and improvement.
  • Design and ship reliability-focused engineering work directly within production codebases, including load balancing, load shedding, instrumentation, scalability, and efficiency improvements.
  • Build and evolve internal infrastructure and developer tooling that enables product engineering teams to independently operate reliable workloads.
  • Partner closely with product development teams to co-design systems and ensure reliability, security, maintainability, and operational readiness are considered throughout development.
  • Develop observability capabilities that make system behavior measurable, understandable, and actionable, using appropriate signals and visualization techniques.
  • Participate in a shared on-call rotation and contribute to effective incident response, investigation, remediation, and prevention of recurring reliability issues.
  • Investigate complex system-level problems across distributed infrastructure, networking, application behavior, and production environments.
  • Improve configuration management and infrastructure practices across diverse systems, reducing unnecessary complexity, errors, and technical debt.
  • Contribute technical perspectives to architectural discussions and help establish engineering practices that support both short-term delivery and long-term scalability.
  • Support systems with demanding workloads, including real-time media, secure customer code execution, advanced networking, and other highly concurrent services.
  • Collaborate with engineering partners on potentially contentious reliability and operational decisions with clarity, pragmatism, and strong technical judgment.
  • Automate repetitive operational processes wherever possible to improve engineering efficiency and reduce manual intervention.

Kenntnisse

Distributed systems design
Observability concepts
Incident response
System troubleshooting
Communication & collaboration
Reliability engineering mindset
Cost-conscious engineering

Tools

Kubernetes
Kafka
ClickHouse

Jobbeschreibung

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer - Reliability, Infrastructure, and Tooling based in Germany.

This is a senior engineering role focused on building reliable, scalable infrastructure for demanding real-time and AI workloads.
You’ll work closely with product engineering teams to design systems that are secure, maintainable, observable, and resilient at global scale.
The role goes beyond traditional operations, combining infrastructure engineering, reliability, tooling, and hands‑on product development.
You’ll tackle complex challenges involving distributed systems, Kubernetes, networking, real-time media, and secure execution environments.
Your work will help engineering teams self‑serve infrastructure and reliability capabilities without creating unnecessary operational bottlenecks.
You’ll also contribute to incident management, on‑call practices, architectural decisions, and long‑term reliability improvements.
This is an opportunity to work remotely with experienced engineers while contributing to infrastructure supporting billions of real‑time interactions.

Accountabilities
  • Ramp up on a complex global architecture involving distributed databases, messaging systems, networking infrastructure, Kubernetes, and other core platform technologies, identifying areas of reliability debt and improvement.

  • Design and ship reliability‑focused engineering work directly within production codebases, including load balancing, load shedding, instrumentation, scalability, and efficiency improvements.

  • Build and evolve internal infrastructure and developer tooling that enables product engineering teams to independently operate reliable workloads.

  • Partner closely with product development teams to co‑design systems and ensure reliability, security, maintainability, and operational readiness are considered throughout development.

  • Develop observability capabilities that make system behavior measurable, understandable, and actionable, using appropriate signals and visualization techniques.

  • Participate in a shared on‑call rotation and contribute to effective incident response, investigation, remediation, and prevention of recurring reliability issues.

  • Investigate complex system‑level problems across distributed infrastructure, networking, application behavior, and production environments.

  • Improve configuration management and infrastructure practices across diverse systems, reducing unnecessary complexity, errors, and technical debt.

  • Contribute technical perspectives to architectural discussions and help establish engineering practices that support both short‑term delivery and long‑term scalability.

  • Support systems with demanding workloads, including real‑time media, secure customer code execution, advanced networking, and other highly concurrent services.

  • Collaborate with engineering partners on potentially contentious reliability and operational decisions with clarity, pragmatism, and strong technical judgment.

  • Automate repetitive operational processes wherever possible to improve engineering efficiency and reduce manual intervention.

Requirements
  • Strong professional experience building and operating non‑trivial production applications, particularly systems involving high concurrency, distributed workloads, or complex control loops.

  • Significant experience with Kubernetes or an equivalent large‑scale container orchestration platform.

  • Strong understanding of Linux internals and networking, with the ability to investigate issues across multiple layers of a production system.

  • Proven experience using observability, monitoring, logging, and related tooling to diagnose difficult production problems.

  • Experience operating large‑scale, globally distributed systems, including the configuration management, reliability challenges, and technical debt that emerge as systems grow.

  • Experience responding to and managing complex production incidents, including identifying root causes and implementing durable corrective actions.

  • Experience operating open‑source infrastructure technologies such as Kafka, ClickHouse, or comparable distributed systems.

  • Strong systems‑thinking skills and an ability to reason about infrastructure in terms of signals, feedback, dependencies, and control mechanisms.

  • Strong communication and collaboration skills, particularly when working with partner engineering teams and navigating competing priorities.

  • A pragmatic approach to engineering that balances immediate delivery needs with long‑term maintainability, reliability, and operational cost.

  • A strong interest in observability, reliability engineering, clean configuration, automation, and reducing operational complexity.

  • Nice to have: experience with data engineering and analytics.

  • Nice to have: experience with global Layer 3 networking.

  • Nice to have: experience operating systems with long‑lived workloads such as real‑time media.

  • Nice to have: experience in Google SRE or another high‑scale reliability engineering environment.

  • Nice to have: experience working with compliance frameworks such as PCI.

Benefits
  • $135,000–$300,000 USD compensation range.

  • Equity participation as part of the overall compensation package.

  • Fully remote work with opportunities for collaboration across a globally distributed organization.

  • Health, dental, and vision benefits.

  • Flexible vacation policy.

  • Opportunity to work on infrastructure supporting large‑scale real‑time and AI applications.

  • Opportunity to contribute to open‑source projects alongside experienced engineers.

  • Exposure to challenging distributed‑systems problems involving real‑time media, secure compute, networking, Kubernetes, and globally distributed infrastructure.

  • Opportunity to influence reliability architecture, engineering practices, and internal developer tooling.

  • Shared on‑call practices designed to keep production experience connected across infrastructure and product engineering teams.

  • Equal opportunity employment and reasonable accommodation support throughout the hiring process.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer / Kubernetes
Senior Site Reliability Engineer / Kubernetes

Jobgether • Deutschland

Remote
EUR 90.000 - 120.000
Senior Site Reliability Engineer (SRE, Compute Node Team)
Senior Site Reliability Engineer (SRE, Compute Node Team)

Jobgether • Deutschland

Vor Ort
EUR 90.000 - 120.000
Competitive pay
Career growth
Flexible work
+2
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether • Deutschland

Vor Ort
EUR 120.000 - 160.000
Competitive compensation
Learning opportunities
Ownership of projects
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

B Capital • Deutschland

Remote
EUR 46.000 - 105.000
Work from anywhere
Flexible paid time off
Mental health support services
+3
Site Reliability Engineering Architect
Site Reliability Engineering Architect

Cavendish Professionals • Berlin

Hybrid
EUR 80.000 - 110.000
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FACT-Finder • Berlin

Hybrid
EUR 110.000 - 150.000
Hybrid work model
Flexible work policy
AI-driven environment
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Meyandy LLC • Berlin

Vor Ort
EUR 90.000 - 130.000
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FactFinder • Berlin

Vor Ort
Confidential
Hybrid work model
Design Engineer
Design Engineer

Jobgether • Deutschland

Remote
EUR 90.000 - 140.000
Remote-first workflow in Europe
Autonomy and impact on product
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FACT-Finder • Pforzheim

Hybrid
EUR 90.000 - 125.000
Hybrid work model