Platform Reliability Lead – AI-native Operations (w/m/d)

Comparus GmbH

Hamburg

Vor Ort

EUR 70.000 - 90.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Eine zielgenaue Bewerbung für diesen Job — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Attractive remuneration
Personalized development
Short decision-making processes

Zusammenfassung

Comparus GmbH in Hamburg is seeking a Platform Reliability Lead to oversee the stable, secure, and cost-effective operation of their platforms. You will be responsible for Operations management, automation, and AIOps while promoting collaboration and accountability. This role offers creative freedom, the chance to build AI-driven operations, and a hybrid work model. The ideal candidate has experience in Platform Operations, SRE, and DevOps, with strong skills in incident management and automation.

Qualifikationen

  • Several years’ experience in Platform Operations, SRE, DevOps, or IT Operations Leadership.
  • In-depth knowledge of modern cloud and platform architectures.
  • Proficiency in SLO/SLA management, incident and risk management.

Aufgaben

  • Overall responsibility for availability, performance, security, and costs.
  • Further development and consolidation of operations, DevOps, and SRE structures.
  • Management of major incidents, including structured internal and external communication.

Kenntnisse

Platform Operations
SRE
DevOps
Incident Management
Automation
AI or AIOps
Decision-Making
Communication

Jobbeschreibung

# Platform Reliability Lead – AI-native Operations (f/m/d)full timehybrid work possibleopen ended## Role descriptionThis role is not a CTO position, nor is it intended as a stepping stone to a C-level role.We deliberately operate without a C-level structure and expect leadership to be demonstrated through accountability and collaboration, not through titles.We are looking for someone who can design operations in such a way that availability, security, costs and scalability work in harmony – even as the platform, product and organisation evolve.## ContextCOMPARUS develops solutions for agile IT transformation and operates its own scalable SaaS and AI platform, TiONA. Depending on the client’s specific needs and regulatory requirements, our platform can be hosted in our own cloud, in public cloud environments, or on-premises at the client’s site.We deliberately combine professional services and product development: services fund innovation, products scale impact.For this to work, we need operations that do not merely manage, but enable sustainable growth.## Your missionAs Platform Reliability Lead, you will have end-to-end responsibility for the stable, secure and cost-effective operation of our platforms. This responsibility is not fulfilled through a supervisory role,but through active involvement in operations and system management.Your goal is clear:Operations must not be a bottleneck – they must enable scaling without causing complexity and costs to spiral out of control.You combine traditional operational excellence with automation, AI and AIOps to ensure reliability not through manual intervention, but through a systematic approach.## Your area of responsibility### System-wide responsibility* Overall responsibility for availability, performance, security and costs* Establishing and managing SLOs, operational KPIs and risk assessments* Defining clear decision-making processes for release, change and go/no-go decisions* Ensuring genuine 24/7 operational capability across relevant environments### DevOps, SRE & Incident Management* Further development and consolidation of operations, DevOps and SRE structures* Establishment and leadership of a small, high-impact operations/DevOps team* Establishment of clear incident, problem and change processes* Management of major incidents, including structured internal and external communication### AI, Automation & AIOps* Identification and prioritisation of genuine automation levers along the Ops value chain* Introduction and utilisation of AIOps approaches, including: + Intelligent alerting & alert deduplication + Anomaly detection & predictive monitoring + Incident summaries & runbook automation + Capacity and cost forecasting* Goal: fewer manual interventions, greater stability through systems### Costs, compliance & scaling* Responsibility for cloud and platform cost management (FinOps)* Ensuring data protection, information security and auditability* Preparing the platform for growth, new customers and new markets* Setting up transparent dashboards for technical, business and management## Expected impactAfter 30 days* You have a clear picture of the platform’s status, risks and operational bottlenecks* You can clearly prioritise what is currently holding back scaling – whether technical, organisational or financialAfter 60 days* Key SLOs are defined, measurable and embedded in day-to-day operations* Incident and on-call structures ensure reliable responsiveness without burning out teams* Initial automations noticeably reduce manual ops workAfter 90 days* Operations are no longer a bottleneck for product or customer scaling* AIOps approaches deliver tangible value (fewer alerts, faster analyses, better forecasts)* Costs, risks and operational capability are transparent and manageable6–12-month outlook* The platform scales without operational overhead increasing linearly* Operations are viewed internally as an enabler for growth, not as a controlling body* Decisions regarding architecture, releases and scaling are based on data rather than gut feeling## What we are looking for* Several years’ experience in Platform Operations, SRE, DevOps or IT Operations Leadership* In-depth knowledge of modern cloud and platform architectures* Experience in building and leading cross-functional teams* Proficiency in SLO/SLA management, incident and risk management* Proven experience with automation and AI or AIOps approaches* Strong decision-making skills, clear communication and a structured approach to work* Entrepreneurial mindset with a focus on quality, efficiency and scalabilityImportant:For us, leadership means leading by example. You will consistently take a hands-on approach – particularly when it comes to critical operational decisions, incident analysis, system design, automation and the further development of reliability and AIOps structures.This role is not purely a management role.## What we offer* A strategically central role with genuine scope for creativity* Work on scaling AI and SaaS platforms* The opportunity to establish AI-driven operations from the ground up* Short decision-making processes, no title-based hierarchy, high level of personal responsibility* International environment with a hybrid working model* Attractive remuneration and personalised development## Cultural fitOur focus isn’t always constant – but our sense of responsibility is.If you can handle shifting priorities whilst still delivering reliable results, you’ll find plenty of creative freedom here.If, however, you require long-term stability and unchanging conditions, this role is probably not the right fit for you.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Head of Engineering – AI-native Platform (w/m/d)
Head of Engineering – AI-native Platform (w/m/d)

Comparus GmbH • Hamburg

Hybrid
EUR 140.000 - 190.000
Remote in Germany or hybrid (Hamburg)
International environment
Attractive remuneration
Platform Product & Growth Lead – AI Platform (TiONA) (w/m/d)
Platform Product & Growth Lead – AI Platform (TiONA) (w/m/d)

Comparus GmbH • Hamburg

Hybrid
EUR 90.000 - 130.000
Attractive remuneration
Close collaboration with teams
Short decision-making processes
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Meyandy LLC • Berlin

Hybrid
EUR 90.000 - 130.000
Sales & Growth Lead – AI-native Platform (w/m/d)
Sales & Growth Lead – AI-native Platform (w/m/d)

Comparus GmbH • Hamburg

Remote
EUR 80.000 - 100.000
Attraktive Vergütung
Flexible Arbeitszeiten
Möglichkeiten zur Weiterbildung
Senior Site Reliability Engineer
Senior Site Reliability Engineer

CloudFactory • Berlin

Vor Ort
EUR 90.000 - 130.000
Senior Cloud Engineer, AI Platform SRE
Senior Cloud Engineer, AI Platform SRE

xdesign • Deutschland

Hybrid
EUR 90.000 - 130.000
Flexible working
Office hubs (UK, Europe)
Career growth & training
Senior/Staff Platform Engineer (m/f/x)
Senior/Staff Platform Engineer (m/f/x)

DUDE CHEM • Berlin

Hybrid
EUR 120.000 - 180.000
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

FACT-Finder • Pforzheim

Vor Ort
EUR 110.000 - 170.000
Competitive salary
Modern equipment
Learning budget
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jackalope Digital LLC • Berlin

Hybrid
EUR 90.000 - 140.000
Team Lead - Site Reliability Engineering (all genders)
Team Lead - Site Reliability Engineering (all genders)

Meyandy LLC • Berlin

Vor Ort
EUR 110.000 - 150.000
Competitive salary
Learning budget
Hybrid work model
+1