Observability Platform Engineer — Neocloud

Mirantis

Warszawa

On-site

PLN 190,000 - 280,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mirantis is expanding its Neocloud service to manage large-scale infrastructure with high SLA. We are seeking an Observability Platform Engineer to design and build the monitoring, logging, tracing, and alerting stack that enables rapid incident detection and resolution across a globally distributed environment.

This hands-on role focuses on turning data visibility into actionable signals, building scalable telemetry pipelines, and partnering with operations teams to align dashboards and alerts

Responsibilities

  • Design, build, and operate observability platform components — metrics, logging, distributed tracing, and alerting — for large-scale infrastructure environments
  • Build telemetry pipelines capable of handling high cardinality, high volume data from large fleets of infrastructure, with an eye on cost, retention, and query performance
  • Define and implement SLO/SLI frameworks and alerting strategies that reduce noise and surface real signal to on-call engineers
  • Partner closely with service delivery and operations teams to understand what they need to see during an incident, and build for that — not just for dashboards nobody opens
  • Integrate observability tooling with incident management workflows, including root-cause analysis support and post-incident review data
  • Continuously improve detection speed and reduce mean-time-to-detect (MTTD) and mean-time-to-resolve (MTTR) across the platform
  • Own the reliability, scalability, and security of the observability stack itself — it needs to be up when everything else is on fire
  • Document architecture, runbooks, and operational practices so the platform is maintainable beyond you

Job description

About the Role

Mirantis is building out our Neocloud service offering — managing large-scale infrastructure to a high SLA for customers running demanding compute workloads. As that offering scales, so does the volume and complexity of telemetry we need to collect, correlate, and act on. We're looking for an Observability Platform Engineer to design and build the monitoring, logging, tracing, and alerting platform that our operations teams depend on to detect and resolve incidents fast — at large scale, across a globally distributed environment. This is a hands-on, build-it role. You'll be the person who turns "we have no visibility into this" into a platform that surfaces the right signal at the right time, and turns "we found out from the customer" into "we caught it before they noticed."

Key Responsibilities
  • Design, build, and operate observability platform components — metrics, logging, distributed tracing, and alerting — for large-scale infrastructure environments
  • Build telemetry pipelines capable of handling high cardinality, high volume data from large fleets of infrastructure, with an eye on cost, retention, and query performance
  • Define and implement SLO/SLI frameworks and alerting strategies that reduce noise and surface real signal to on-call engineers
  • Partner closely with service delivery and operations teams to understand what they need to see during an incident, and build for that — not just for dashboards nobody opens
  • Integrate observability tooling with incident management workflows, including root-cause analysis support and post-incident review data
  • Continuously improve detection speed and reduce mean-time-to-detect (MTTD) and mean-time-to-resolve (MTTR) across the platform
  • Contribute to the roadmap for AI-assisted operations tooling (e.g., automated triage, anomaly detection, engineer-assist tooling) as it m rovide…? Actually continues— ...
  • Own the reliability, scalability, and security of the observability stack itself — it needs to be up when everything else is on fire
  • Document architecture, runbooks, and operational practices so the platform is maintainable beyond you
What Success Looks Like
  • Operations teams can diagnose incidents faster because the right data is surfaced automatically, not hunted for manually
  • Alert volume is high-signal, low-noise — engineers trust what fires
  • The observability platform scales cleanly as infrastructure footprint grows, without cost or performance surprises
  • Reduced MTTD/MTTR trends, tracked and demonstrable over time
  • A platform other engineers actually want to build on, not work around
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Observability Platform Engineer — Scale & Incident Readiness
Observability Platform Engineer — Scale & Incident Readiness

Mirantis • Warszawa

On-site
PLN 190,000 - 280,000
Senior MLOps & Observability Engineer | Kubernetes Cloud
Senior MLOps & Observability Engineer | Kubernetes Cloud

CloudFerro Sp. z o • Warszawa

On-site
PLN 180,000 - 270,000
Senior Software Engineer-MLOps & Observability
Senior Software Engineer-MLOps & Observability

CloudFerro Sp. z o • Warszawa

On-site
PLN 180,000 - 270,000
Medical care
Multisport
Life insurance
Remote Observability Platform Engineer for Scalable Systems
Remote Observability Platform Engineer for Scalable Systems

Whatnot Inc. • Kraków

Hybrid
PLN 240,000 - 360,000
Remote Observability Engineer - Scale & Reliability
Remote Observability Engineer - Scale & Reliability

Whatnot • Kraków

Hybrid
PLN 520,000 - 580,000
Remote Cloud Infrastructure Service Delivery Manager
Remote Cloud Infrastructure Service Delivery Manager

Mirantis • Poznań

Remote
PLN 180,000 - 240,000
Professional development
Conference attendance
Company events
+1
Observability & Tooling Manager
Observability & Tooling Manager

Planet • Warszawa

Hybrid
PLN 260,000 - 380,000
Hybrid work model
Software Engineer, Observability
Software Engineer, Observability

Whatnot Inc. • Kraków

Hybrid
PLN 240,000 - 360,000
Lead DevOps Engineer — Observability & CI/CD Champion
Lead DevOps Engineer — Observability & CI/CD Champion

Duco Technology Ltd • Wrocław

On-site
PLN 180,000 - 240,000
Observability & Tooling Leader, Cloud-Native Reliability
Observability & Tooling Leader, Cloud-Native Reliability

Planet • Warszawa

Hybrid
PLN 260,000 - 380,000
Hybrid work model