Principal Engineer, CSRE Provisioning

Jobtailor

Deutschland

Remote

EUR 100.000 - 130.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Erhalte eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Jobtailor is seeking a Senior Platform Reliability Engineer to lead multi-quarter lifecycle work across the CSRE hosted estate. You will set technical direction for platform life, guide high-risk initiatives, and drive robust reliability and observability practices.

In this role you will act as escalation point for incidents, participate in architecture reviews with Platform Engineering, and mentor engineers while advancing monitoring, DR validation, and on-call readiness across complex

Qualifikationen

  • Extensive practical experience operating and evolving production infrastructure at scale.
  • Expert understanding of SRE principles, including SLIs, SLOs, and error budgets.
  • Deep expertise designing and troubleshooting distributed systems.
  • Experience with on-premises data centers and cloud-native environments.
  • Ability to participate in on-call rotations and mentor other engineers.

Aufgaben

  • Set technical direction and a multi-quarter roadmap for platform lifecycle work across the CSRE hosted estate.
  • Lead the hardest and highest-risk platform initiatives, including platform-side execution.
  • Define and evolve Provisioning’s supported operating model and standards.
  • Act as a technical escalation point for production incidents and drive durable structural fixes.
  • Represent Provisioning in architecture and design reviews with Platform Engineering, application teams, and senior leadership.
  • Build and evolve reliability, disaster recovery, and observability practices across the estate.
  • Lead service acceptance for complex or high-risk new platforms.
  • Identify infrastructure and processes that can be retired before advocating for new ones.
  • Participate in the team’s on-call rotation and escalates complex incidents.
  • Mentor and develop senior and staff-level engineers through pairing, design review, and career coaching.

Kenntnisse

Production Infrastructure
Distributed Systems
SRE Principles
Infrastructure-as-Code
Incident Analysis
Automation
Governance
On-Prem & Cloud
Jira

Tools

Jira
Cloud Native Environments
On-Prem Infra

Jobbeschreibung

• Set technical direction and a multi-quarter roadmap for platform lifecycle work across the CSRE hosted estate
• Lead the hardest and highest-risk platform initiatives, including platform-side execution
• Define and evolve Provisioning’s supported operating model and standards
• Act as a technical escalation point for production incidents and drive durable structural fixes
• Represent Provisioning in architecture and design reviews with Platform Engineering, application teams, and senior leadership
• Build and evolve reliability, disaster recovery, and observability practices across the estate
• Lead service acceptance for complex or high-risk new platforms
• Identify infrastructure and processes that can be retired before advocating for new ones
• Participate in the team’s on-call rotation and escalates complex incidents
• Mentor and develop senior and staff-level engineers through pairing, design review, and career coaching
• Contribute to sprint planning and PI-cycle delivery on the PROV Jira board
• Feed operational insights back into CSRE Platform tooling and Provisioning standards

Requirements
  • Extensive practical experience operating and evolving production infrastructure at scale
  • Experience with complex, high-risk migrations and decommissioning
  • Deep expertise designing and troubleshooting distributed systems
  • Expert understanding of SRE principles, including SLIs, SLOs, and error budgets
  • Extensive experience with on-premises data center infrastructure and cloud-native environments
  • Experience with governance, cost trade-offs, and vendor evaluation at scale
  • Track record improving observability and alerting across many platforms
  • Deep experience with infrastructure-as-code and automation
  • Excellent software engineering fundamentals
  • Extensive experience with production readiness, DR validation, and controlled failure testing
  • Expert-level incident analysis skills
  • Practical experience directing where LLM and AI-assisted tooling helps engineering work
  • Excellent written and verbal communication
  • Ability to participate in an on-call rotation
  • Ability to mentor senior and staff-level engineers
Core Competencies

Demonstrates extensive experience in operating and evolving production infrastructure at scale, with a strong focus on incident analysis, observability, and disaster recovery practices. Capable of mentoring engineers and leading high-risk platform initiatives while ensuring adherence to SRE principles.

Highest-signal resume keywords
  • Production Infrastructure Management
  • Distributed Systems Design
  • SRE Principles Expertise
  • Infrastructure-as-Code
  • Incident Analysis
Hard Skills
  • Production Readiness
  • Disaster Recovery Validation
  • Error Budgets
  • Automation
  • Observability Improvement
  • Software Engineering Fundamentals
  • High-Risk Migrations
  • Decommissioning
  • Governance
  • Cost Trade-offs
Soft Skills
  • Excellent Written Communication
  • Excellent Verbal Communication
  • Mentoring
Industry Keywords
  • Platform Lifecycle
  • Service Acceptance
  • Technical Direction
  • Production Incidents
  • Operational Insights
Tools & Technologies
  • Cloud-Native Environments
  • On-Premises Data Center Infrastructure
  • Jira
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Mid-Level SRE
Mid-Level SRE

Jobtailor • Deutschland

Hybrid
EUR 70.000 - 110.000
Director – Platform & Infrastructure
Director – Platform & Infrastructure

Jobtailor • Deutschland

Remote
EUR 120.000 - 180.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Meyandy LLC • Berlin

Hybrid
EUR 90.000 - 130.000
Infrastructure Engineer
Infrastructure Engineer

Jobtailor • Deutschland

Remote
EUR 70.000 - 110.000
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

delinea • Deutschland

Hybrid
EUR 120.000 - 180.000
Healthcare insurance
Pension plan
Life insurance
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

CloudFactory Limited • Berlin

Vor Ort
EUR 90.000 - 130.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

CloudFactory • Berlin

Vor Ort
EUR 90.000 - 130.000
Engineer II – Platform
Engineer II – Platform

Jobtailor • Deutschland

Remote
EUR 70.000 - 110.000
Site Reliability Engineer
Site Reliability Engineer

Apprize Technology Solutions • Deutschland

Vor Ort
EUR 70.000 - 90.000
Principal Platform Engineer, AI Engineering
Principal Platform Engineer, AI Engineering

RxSense • Deutschland

Hybrid
EUR 120.000 - 190.000