Mid-Level SRE

Jobtailor

Deutschland

Hybrid

EUR 70.000 - 110.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Jobtailor is seeking a seasoned SRE/DevOps professional to join our production reliability team in Germany. You will prevent incidents, build and refine our observability platform, and lead postmortems to drive actionable improvements.

You will work with AWS and on-prem environments, emphasize automation and cost predictability, and collaborate with developers to elevate reliability and performance across services.

Qualifikationen

  • Solid experience troubleshooting production environments and distributed systems.
  • Experience with AWS (EC2, networking, load balancing, IAM) and on-prem environments a plus.
  • Knowledge of Docker/Docker Compose for production.
  • Experience with observability/monitoring tools (Grafana, Prometheus, Datadog, SigNoz, OpenTelemetry).
  • Understanding of SRE practices: SLI/SLO, incident management, and postmortems.
  • Strong Linux, networking, and protocol knowledge (HTTP, TCP/IP, DNS).
  • Strong communication, autonomy, and resilience during critical incidents.
  • Experience with automation (Python, Bash, Terraform, Ansible) a plus.
  • Familiarity with DevOps: CI/CD and IaC.

Aufgaben

  • Prevent production incidents by identifying operational risks and noisy alerts.
  • Contribute to building the observability platform (logs, metrics, tracing) and define SLIs/SLOs.
  • Lead root cause analyses and postmortems, tracking action plans to completion.
  • Resolve incidents in AWS and on-prem production environments.
  • Identify FinOps opportunities and contribute to cloud cost predictability.
  • Collaborate on improving reliability and performance with development teams.
  • Support scalability initiatives and automation-focused infrastructure.
  • Contribute to SRE maturity through documentation and incident culture.

Jobbeschreibung

  • Prevent production incidents by identifying operational risks, points of failure, and noisy alerts before they become problems.
  • Participate in building the observability platform (logs, metrics, and tracing), contributing to the definition of processes, SLIs, SLOs, and error budgets.
  • Conduct root cause analyses and lead postmortems, documenting lessons learned and tracking action plans through to completion.
  • Resolve critical incidents through troubleshooting in AWS and on-premises production environments.
  • Identify FinOps opportunities and contribute to cloud cost predictability.
  • Collaborate with the development team on the continuous improvement of application reliability and performance.
  • Support scalability initiatives and the creation of new infrastructure, with a focus on automation.
  • Contribute to the team’s SRE maturity through practices, documentation, and incident management culture.
Requirements
  • Solid experience troubleshooting production environments and distributed systems.
  • Production experience with AWS (EC2, networking, load balancing, IAM); experience with on-premises environments is a plus.
  • Knowledge of Docker/Docker Compose, including running containers in production.
  • Experience with observability and monitoring tools (e.g., Grafana, Prometheus, Datadog, SigNoz, or similar) and OpenTelemetry (logs, metrics, and tracing).
  • Understanding of SRE practices: SLI/SLO, error budgets, incident management and resolution, and postmortems.
  • Strong knowledge of Linux, networking, and protocols (HTTP, TCP/IP, DNS).
  • Strong communication skills, autonomy, and resilience when working during critical incidents, including occasional direct interaction with customers.
  • Experience with automation (Python, Bash, Terraform, Ansible, or similar) is a plus.
  • Familiarity with DevOps practices, including CI/CD and infrastructure as code (IaC), is a plus.
Core Competencies

Demonstrates expertise in troubleshooting production environments and distributed systems, with a strong focus on AWS, observability tools, and SRE practices. Capable of driving incident management and resolution while contributing to infrastructure automation and cost predictability.

Highest-signal resume keywords
  • AWS Production Experience
  • Observability Tools Proficiency
  • SRE Practices Knowledge
  • Linux Networking Expertise
  • Automation Skills
Hard Skills
  • Troubleshooting Production Environments
  • Distributed Systems
  • Docker/Docker Compose
  • OpenTelemetry
  • SLI/SLO Understanding
  • Incident Management
  • Python
  • Bash
  • Terraform
  • Ansible
Soft Skills
  • Strong Communication Skills
  • Autonomy
  • Resilience
Industry Keywords
  • FinOps
  • Cloud Cost Predictability
  • Infrastructure as Code (IaC)
  • CI/CD
  • Incident Management Culture
Tools & Technologies
  • AWS (EC2, Networking, Load Balancing, IAM)
  • Grafana
  • Prometheus
  • Datadog
  • SigNoz
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Principal Engineer, CSRE Provisioning
Principal Engineer, CSRE Provisioning

Jobtailor • Deutschland

Remote
EUR 100.000 - 130.000
Site Reliability Engineer
Site Reliability Engineer

Apprize Technology Solutions • Deutschland

Vor Ort
EUR 70.000 - 90.000
SRE Engineer (Fully Remote)
SRE Engineer (Fully Remote)

Not Disclosed • Deutschland

Hybrid
EUR 80.000 - 120.000
Meal allowance
Home office allowance
Medical insurance
+2
Infrastructure Engineer
Infrastructure Engineer

Jobtailor • Deutschland

Remote
EUR 70.000 - 110.000
Senior Site Reliability Expert
Senior Site Reliability Expert

Upserve • Deutschland

Hybrid
EUR 80.000 - 110.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Meyandy LLC • Berlin

Hybrid
EUR 90.000 - 130.000
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

FACT-Finder • Pforzheim

Hybrid
EUR 90.000 - 125.000
Hybrid work model
DevOps/SRE
DevOps/SRE

Embedded Shishya • Deutschland

Remote
EUR 70.000 - 95.000
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

delinea • Deutschland

Hybrid
EUR 120.000 - 180.000
Healthcare insurance
Pension plan
Life insurance
+1
Senior Linux Infrastructure Engineer
Senior Linux Infrastructure Engineer

Tastylive • Deutschland

Hybrid
EUR 121.000 - 164.000