Site Reliability Engineer

Jobtailor

Deutschland

Vor Ort

EUR 75.000 - 110.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Jobtailor is seeking a Site Reliability Engineer to own the reliability of .NET/C# services on Windows and participate in on‑call rotations for production trading systems. You will lead incident response and drive RCA actions to prevent recurrence, while building dashboards in Grafana and alerts in Prometheus to enhance visibility and resilience.

You'll instrument services for telemetry, define SLIs/SLOs, and collaborate with developers to improve deployment safety, rollback strategies, and

Qualifikationen

  • 3–5 years of experience in Site Reliability Engineering or related field.
  • Strong experience debugging and supporting .NET/C# applications in production.
  • Hands-on experience with Windows Server environments.
  • Strong PowerShell scripting skills.
  • Experience with Python or Bash.
  • Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools).
  • Experience with modern CI/CD pipelines.
  • Knowledge of deployment strategies, release automation, and rollback mechanisms.
  • Experience working with AWS.
  • Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools.
  • Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms.
  • Practical experience with SLIs & SLOs, Error Budgets, Incident Response, Root Cause Analysis (RCA), and Alert Design.

Aufgaben

  • Own day-to-day reliability of .NET/C# services on Windows.
  • Lead incident response during service disruptions in production.
  • Investigate incidents, perform RCA, and implement preventive actions.
  • Build Grafana dashboards, Prometheus alerts, and health views across apps and infra.
  • Instrument .NET services for telemetry and visibility.
  • Define, implement, and monitor SLIs, SLOs, and error budgets.
  • Troubleshoot across .NET/C# apps, Windows Server, Aurora PostgreSQL, and AWS.
  • Improve deployment safety, release automation, and rollback strategies.
  • Partner with developers to improve operability, resilience, and fault isolation.
  • Automate operational tasks via scripting and IaC.
  • Create and maintain runbooks and incident response procedures.
  • Continuously improve monitoring, alert quality, automation, and platform reliability.

Kenntnisse

Site Reliability Engineering
Debugging .NET/C# Applications
PowerShell Scripting
Grafana and Prometheus
AWS Infrastructure
Terraform
CI/CD Pipelines
Aurora PostgreSQL
Python
Bash
SLIs & SLOs
Incident Response
Root Cause Analysis
Windows Server
Monitoring & Observability

Tools

Windows Server
CI/CD Pipelines
Infrastructure as Code (IaC)
Terraform

Jobbeschreibung

  • Own the day-to-day reliability of our .NET/C# services running on Windows.
  • Participate in the on-call rotation for production trading systems and lead incident response during service disruptions.
  • Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues.
  • Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases.
  • Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact.
  • Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Troubleshoot issues across .NET/C# applications, Windows Server, Aurora PostgreSQL databases, and AWS infrastructure.
  • Improve deployment safety, release automation, and rollback strategies.
  • Partner with developers to improve application operability, resilience, and fault isolation.
  • Automate operational tasks through scripting and infrastructure automation.
  • Create and maintain runbooks, operational documentation, and incident response procedures.
  • Continuously improve monitoring, alert quality, automation, and platform reliability.
Requirements
  • 3–5 years of experience in Site Reliability Engineering or related field
  • Strong experience debugging and supporting .NET/C# applications in production.
  • Hands‑on experience with Windows Server environments.
  • Strong PowerShell scripting skills.
  • Experience with Python or Bash.
  • Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools).
  • Experience with modern CI/CD pipelines.
  • Knowledge of deployment strategies, release automation, and rollback mechanisms.
  • Experience working with AWS.
  • Hands‑on experience with Terraform or other Infrastructure as Code (IaC) tools.
  • Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms.
  • Practical experience with SLIs & SLOs, Error Budgets, Incident Response, Root Cause Analysis (RCA), and Alert Design.
Core Competencies

Demonstrates expertise in Site Reliability Engineering with a focus on .NET/C# applications, Windows Server environments, and AWS infrastructure. Proficient in incident response, monitoring, and automation to enhance system reliability and operational efficiency.

Highest-signal resume keywords
  • Site Reliability Engineering
  • Debugging .NET/C# Applications
  • PowerShell Scripting
  • Grafana and Prometheus
  • AWS Infrastructure
ATS Optimization Keywords
Hard Skills
  • Site Reliability Engineering
  • Debugging .NET/C# Applications
  • PowerShell Scripting
  • Python
  • Bash
  • Grafana
  • Prometheus
  • Terraform
  • Aurora PostgreSQL
  • SLIs and SLOs
Soft Skills
  • Incident Response
  • Root Cause Analysis
Industry Keywords
  • Deployment Strategies
  • Release Automation
  • Rollback Mechanisms
  • Operational Documentation
  • Monitoring and Observability
Tools & Technologies
  • Windows Server
  • CI/CD Pipelines
  • Infrastructure as Code (IaC)
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer, SRE, Backend
Senior Site Reliability Engineer, SRE, Backend

Jobtailor • Berlin

Vor Ort
EUR 90.000 - 130.000
Senior Azure Platform Engineer
Senior Azure Platform Engineer

Jobtailor • Deutschland

Hybrid
EUR 90.000 - 130.000
Senior Full-Stack Developer, .Net, React
Senior Full-Stack Developer, .Net, React

Jobtailor • Deutschland

Remote
EUR 90.000 - 120.000
Infrastructure Engineer, Mid-Senior Level – AWS, Kubernetes
Infrastructure Engineer, Mid-Senior Level – AWS, Kubernetes

Jobtailor • Berlin

Vor Ort
EUR 90.000 - 130.000
Staff Engineer – Founding
Staff Engineer – Founding

Jobtailor • München

Vor Ort
EUR 120.000 - 160.000
Head of DevOps
Head of DevOps

Jobtailor • Deutschland

Vor Ort
EUR 120.000 - 180.000
Software Engineering Manager
Software Engineering Manager

Jobtailor • Deutschland

Vor Ort
EUR 120.000 - 180.000
DevOps Engineer – EU Night Shift
DevOps Engineer – EU Night Shift

Jobtailor • Deutschland

Hybrid
EUR 80.000 - 120.000
Senior .Net Developer
Senior .Net Developer

Jobtailor • Deutschland

Hybrid
EUR 90.000 - 130.000
Site Reliability Engineer
Site Reliability Engineer

Apprize Technology Solutions • Deutschland

Vor Ort
EUR 70.000 - 90.000