Platform Reliability Engineer (Observability & AI Platform)

Newpage Solutions

Polska

Remote

PLN 469,000 - 665,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

People-first culture
Smart collaboration
Work-life balance
Growth opportunities
Competitive compensation

Job summary

Newpage Solutions is seeking an experienced Platform Reliability Engineer to build observability, reliability, and cost-monitoring for enterprise AI platforms. You will implement OpenTelemetry instrumentation, tracing, SLIs/SLOs, and proactive alerting across AWS-based services and applications.

You’ll collaborate with AI platform, cloud infra, and app teams to ensure reliability, operational visibility, and optimized token usage and compute costs. Strong scripting and IaC skills are essential.

Qualifications

  • 7+ years of experience in platform engineering, SRE, DevOps, or production infrastructure operations.
  • Hands-on experience with OpenTelemetry SDKs, Collector configuration, instrumentation, and telemetry pipelines.
  • Strong experience implementing distributed tracing, metrics, logging, dashboards, and alerting.
  • Experience defining and implementing reliability metrics, SLOs, error budgets, and availability monitoring.

Responsibilities

  • Observability & Instrumentation: Design and implement OpenTelemetry-based instrumentation across AI agents, services, APIs, and platform infrastructure.
  • AI Agent Monitoring: Implement end-to-end tracing for agent execution, LLM interactions, tool calls, multi-agent workflows, and external API integrations.
  • Platform Reliability: Define and implement Service Level Indicators (SLIs), SLOs, availability monitoring, and reliability dashboards across AI platform components.
  • Alerting & Incident Management: Configure proactive alerts for service degradation, execution failures, latency spikes, and resource utilization.
  • Cost Observability: Build monitoring and reporting capabilities for LLM token consumption, model API costs, AWS compute utilization, and agent execution costs.
  • Observability Integration: Integrate telemetry from AI platforms, AWS services, and applications into centralized observability solutions.
  • Automation & Continuous Improvement: Automate observability configuration, dashboard provisioning, and alert management using IaC.

Skills

OpenTelemetry
Distributed tracing
Telemetry pipelines
Python scripting
Terraform / IaC
Cost monitoring
AWS observability

Education

7+ years Platform Engineering / SRE / DevOps

Tools

Prometheus
Grafana
AWS CloudWatch
EKS
OpenTelemetry SDKs

Job description

About Newpage Solutions

Newpage Solutions is a global digital health innovation company helping people live longer, healthier lives. We partner with life sciences organisations which include, pharmaceutical, biotech and healthcare leaders, to build transformative AI and data driven technologies addressing real-world health challenges.

Location / Contract

LOCATION: Remote | Type: Contract

Your Mission

We are looking for an experienced Platform Reliability Engineer to build and manage observability, reliability, and cost-monitoring capabilities for enterprise AI and agent-based platforms. The role focuses on implementing OpenTelemetry instrumentation, distributed tracing, Service Level Objectives (SLOs), proactive alerting, and monitoring of LLM token consumption and cloud compute costs. The engineer will work closely with AI platform, cloud infrastructure, and application engineering teams to ensure platform reliability, operational visibility, and cost efficiency.

What You Ll Do
  • Observability & Instrumentation: Design and implement OpenTelemetry-based instrumentation across AI agents, services, APIs, and platform infrastructure. Establish centralized logging, metrics, and distributed tracing.
  • AI Agent Monitoring: Implement end-to-end tracing for agent execution, LLM interactions, tool calls, multi-agent workflows, and external API integrations. Monitor execution failures, latency, token usage, and performance.
  • Platform Reliability: Define and implement Service Level Indicators (SLIs), SLOs, availability monitoring, and reliability dashboards across AI platform components.
  • Alerting & Incident Management: Configure proactive alerts for service degradation, execution failures, latency spikes, and resource utilization. Support incident investigation, root cause analysis, and operational runbooks.
  • Cost Observability: Build monitoring and reporting capabilities for LLM token consumption, model API costs, AWS compute utilization, and agent execution costs. Implement cost attribution, budget thresholds, and anomaly alerts.
  • Observability Integration: Integrate telemetry from AI platforms, AWS services, and applications into centralized observability solutions. Enable consistent monitoring and reporting across development, staging, and production environments.
  • Automation & Continuous Improvement: Automate observability configuration, dashboard provisioning, and alert management using Infrastructure as Code (IaC). Identify opportunities to improve platform performance, reliability, and operational efficiency.
What You Bring
  • 7+ years of relevant experience in platform engineering, SRE, DevOps, or production infrastructure operations.
  • Hands-on experience with OpenTelemetry SDKs, Collector configuration, instrumentation, and telemetry pipelines.
  • Strong experience implementing distributed tracing, metrics, logging, dashboards, and alerting.
  • Experience defining and implementing reliability metrics, SLOs, error budgets, and availability monitoring.
  • Experience monitoring AWS infrastructure and services, including CloudWatch, EKS, and containerized workloads.
  • Hands-on experience with Prometheus, Grafana, or equivalent enterprise observability platforms.
  • Experience monitoring and analysing cloud compute costs, resource consumption, and usage trends.
  • Proficiency in Python or another scripting language, with experience in Terraform or equivalent IaC tools.
  • Strong production troubleshooting, incident management, and root cause analysis skills.
Nice To Have Skills
  • Experience implementing observability for LLM applications, AI agents, and multi-agent orchestration.
  • Familiarity with AWS Bedrock, AgentCore, or similar AI runtime platforms.
  • Experience with Langfuse, Lang Smith, or similar LLM observability platforms.
  • Knowledge of LLM token accounting, model pricing, inference latency, and cost attribution.
  • Experience implementing OpenTelemetry GenAI semantic conventions and tracing agent-to-agent interactions.
  • Familiarity with Kubernetes, Docker, GitOps, and CI/CD pipelines.
  • Experience with FinOps practices, cost optimization, and automated budget alerting.
  • Exposure to Grafana Tempo, Loki, Jaeger, or similar observability tools.
Expected Deliverables
  • OpenTelemetry instrumentation and end-to-end distributed tracing across AI platform components.
  • Centralized dashboards covering platform health, service performance, agent execution, and LLM usage.
  • Defined SLIs, SLOs, error budgets, and proactive alerting for critical platform services.
  • Token and compute cost dashboards with usage attribution, budget monitoring, and anomaly detection.
  • Operational runbooks, incident investigation workflows, and automated observability configurations.
Preferred Skills

An experienced Platform Reliability Engineer with strong hands‑on expertise in observability, cloud infrastructure, and production reliability. The candidate should be comfortable working across AWS‑based AI platforms, implementing standardized telemetry, and translating technical operational data into actionable reliability and cost insights.

What We Offer

At Newpage, we re building a company that works smart and grows with agility, where driven individuals come together to do work that matters. We offer:

  • A people-first culture - Supportive peers, open communication and a strong sense of belonging
  • Smart, purposeful collaboration - Work with talented colleagues to create technologies that solve meaningful business challenges
  • Balance that lasts - We respect your time and support a healthy integration of work and life
  • Room to grow - Opportunities for learning, leadership and career development, shaped around you
  • Meaningful rewards - Competitive compensation that recognises both contribution and potential
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AWS_Cloud_Platform_Engineer_AgentCore_Deployment
AWS_Cloud_Platform_Engineer_AgentCore_Deployment

Newpage Solutions • Polska

Remote
PLN 548,000 - 743,000
Remote Platform Reliability Engineer - AI & Observability
Remote Platform Reliability Engineer - AI & Observability

Newpage Solutions • Polska

Remote
PLN 469,000 - 665,000
People-first culture
Smart collaboration
Work-life balance
+2
Lead FDE Production & Consulting
Lead FDE Production & Consulting

Newpage Solutions • Poland

On-site
PLN 180,000 - 300,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Luxoft • Poland

On-site
PLN 180,000 - 260,000
Private Medical & Dental care
Life Insurance covered
Internal Mobility program
Senior AI Platform Engineer
Senior AI Platform Engineer

EPAM Systems • Poland

On-site
PLN 180,000 - 260,000
Hybrid by design
Remote work within Poland
Relocation opportunities
+7
Senior Software Engineer-MLOps & Observability
Senior Software Engineer-MLOps & Observability

CloudFerro Sp. z o • Warszawa

On-site
PLN 180,000 - 270,000
Medical care
Multisport
Life insurance
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hard Rock Digital • Polska

Hybrid
PLN 260,000 - 380,000
AI Engineer
AI Engineer

Omnitrace • Warszawa

On-site
PLN 180,000 - 240,000
Senior MLOps & Observability Engineer | Kubernetes Cloud
Senior MLOps & Observability Engineer | Kubernetes Cloud

CloudFerro Sp. z o • Warszawa

On-site
PLN 180,000 - 270,000
Platform / DevOps Engineer
Platform / DevOps Engineer

Lekta • Kraków

Hybrid
PLN 180,000 - 260,000