Senior Observability Platform Engineer (AI/LLM)

Moderna

Warszawa

Hybrid

PLN 300,000 - 520,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Healthcare
Well-being programs
Family building benefits
Generous paid time off
Professional development
Location-specific perks

Job summary

Moderna Warsaw is expanding its international business services hub and inviting senior engineers to lead the enterprise observability platform. You will own governance, drive architecture standards, and build scalable, cost-efficient observability across applications, databases, hosts, containers, and cloud services.

You will implement AI observability and automation using Python, Terraform, and Langfuse, with a focus on reliability and measurable MTTR improvements.

Qualifications

  • 7+ years of experience in site reliability engineering, observability engineering, platform engineering, or related technical disciplines.
  • Strong hands-on experience designing, implementing, and operating modern observability platforms.
  • Experience with observability technologies such as OpenTelemetry, Prometheus, Grafana, VictoriaMetrics, Elastic, Datadog, Dynatrace, or similar solutions.
  • Strong understanding of metrics, logs, traces, telemetry pipelines, and SLO/SLI frameworks.
  • Experience supporting applications, infrastructure, containers, cloud services, and distributed systems.
  • Experience integrating observability platforms with incident management and operational workflows.
  • Hands-on experience with automation and infrastructure-as-code technologies such as Python, Terraform, Ansible, or Bash.
  • Experience supporting cloud-native and hybrid environments (AWS and/or Azure).
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Strong communication and stakeholder management skills.

Responsibilities

  • Own the enterprise observability platform, driving cost optimization, capacity planning, telemetry governance, and sustainable platform growth.
  • Manage and evolve Moderna's observability platform using technologies including OpenTelemetry, Prometheus, Grafana, VictoriaMetrics, and other modern observability solutions.
  • Lead platform governance, agent lifecycle management, architecture standards, roadmap development, and platform best practices.
  • Collaborate with technology vendors and open-source communities to influence product roadmaps and maximize platform value.
  • Design and build scalable, resilient, and cost-efficient observability architectures supporting applications, databases, hosts, containers, cloud services, networks, distributed systems, and AI/LLM-based workloads.
  • Develop and optimize telemetry pipelines for metrics, traces, and logs across hybrid cloud and on-premises environments.
  • Establish enterprise standards for monitoring, alerting, SLOs, SLIs, and proactive incident detection.
  • Enable self-service observability capabilities that accelerate troubleshooting, operational visibility, and platform reliability.
  • Design and implement observability capabilities for AI agents, LLM-powered applications, and agentic workflows, including monitoring prompts, responses, execution flows, latency, errors, taken consumption, and operational costs.
  • Define enterprise standards for AI observability, monitoring model performance, user interactions, reliability, and cost efficiency.
  • Implement AI observability solutions using Langfuse or similar technologies, enabling prompt analytics, optimization, intelligent alerting, debugging, and failure pattern detection.
  • Leverage AI and LLM technologies to support anomaly detection, operational insights, and root cause analysis.
  • Lead the enterprise logging strategy as a core pillar of the observability platform.
  • Design, build, and scale cost-efficient logging solutions using Grafana Loki and other modern open-source technologies.
  • Define enterprise logging standards, including centralized log ingestion, parsing, querying, retention policies, storage optimization, and query performance.
  • Partner with Security teams to ensure compliance, audit readiness, and forensic capabilities.
  • Integrate observability capabilities with incident management platforms such as PagerDuty to improve operational responsiveness.
  • Optimize on-call processes by ensuring alerts are meaningful, actionable, and routed effectively while supporting rapid incident resolution.
  • Provide real-time telemetry during incidents and contribute to root cause analysis activities.
  • Develop automation using Python, Terraform, Ansible, CI/CD pipelines, and infrastructure-as-code practices.
  • Implement self-healing capabilities and automated remediation to improve operational resilience.
  • Integrate the observability platform with enterprise technologies including ServiceNow, Jira, and other operational systems.
  • Develop dashboards and executive reporting that provide visibility into platform adoption, telemetry coverage, MTTA, MTTR, alert quality, reliability, operational performance, and cost efficiency.
  • Produce documentation, runbooks, knowledge articles, and technical training to support platform adoption and engineering consistency.
  • Participate in post-incident reviews and drive continuous improvement initiatives that strengthen operational excellence and foster a culture of observability and data-driven decision-making.

Skills

SRE
Observability
Platform engineering
Automation
Python
Terraform
Ansible
Cloud (AWS/Azure)
OpenTelemetry
Prometheus
Grafana
ID/Logging

Tools

OpenTelemetry
Prometheus
Grafana
VictoriaMetrics
Elastic
Datadog
Dynatrace
Terraform
Ansible

Job description

Moderna Warsaw is expanding its international business services hub and inviting senior engineers to lead the enterprise observability platform. You will own governance, drive architecture standards, and build scalable, cost-efficient observability across applications, databases, hosts, containers, and cloud services.

You will implement AI observability and automation using Python, Terraform, and Langfuse, with a focus on reliability and measurable MTTR improvements.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior DevSecOps Platform Engineer
Senior DevSecOps Platform Engineer

Moderna, Inc. • Poland

On-site
PLN 180,000 - 320,000
Healthcare
Well‑being programs
Generous paid time off
Platform Engineer: AI-Driven Dev Tooling & AWS
Platform Engineer: AI-Driven Dev Tooling & AWS

Moderna Therapeutics • Warszawa

Hybrid
PLN 180,000 - 240,000
Competitive healthcare
Well-being resources
Family building benefits
+1
Senior Full-Stack Engineer: AI-Assisted Observability & SLOs
Senior Full-Stack Engineer: AI-Assisted Observability & SLOs

Luxoft Poland • Poland

On-site
PLN 180,000 - 240,000
Private Medical & Dental care
Life Insurance
MyBenefit program
+1
Senior Backend Engineer - Observability, Telemetry & AI
Senior Backend Engineer - Observability, Telemetry & AI

Luxoft Poland • Poland

On-site
PLN 120,000 - 180,000
Private Medical & Dental care
Life Insurance
MyBenefit program
+1
AI-Driven Site Reliability Engineer for Cloud Observability
AI-Driven Site Reliability Engineer for Cloud Observability

Luxoft Poland • Poland

On-site
PLN 180,000 - 240,000
Senior Full-Stack Engineer: AI-Driven Observability
Senior Full-Stack Engineer: AI-Driven Observability

Luxoft Germany • Polska

On-site
PLN 180,000 - 240,000
Senior AI-Driven Observability & SRE Engineer
Senior AI-Driven Observability & SRE Engineer

Luxoft Poland • Poland

On-site
PLN 180,000 - 260,000
Private Medical
Dental care
Life Insurance
+1
Platform Engineer: Cloud Infra & DevTools
Platform Engineer: Cloud Infra & DevTools

Moderna • Warszawa

Hybrid
PLN 300,000 - 420,000
Healthcare
Well-being resources
Family building benefits
+3
Senior Observability Engineer: Scalable Telemetry for AI
Senior Observability Engineer: Scalable Telemetry for AI

Coreweave • Wrocław

On-site
PLN 260,000 - 420,000
Living Wage
AI Infra Reliability & Observability Engineer
AI Infra Reliability & Observability Engineer

Margo Group • Warszawa

Hybrid
PLN 180,000 - 260,000