Remote Platform Reliability Engineer - AI & Observability

Newpage Solutions

Polska

Remote

PLN 469,000 - 665,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

People-first culture
Smart collaboration
Work-life balance
Growth opportunities
Competitive compensation

Job summary

Newpage Solutions is seeking an experienced Platform Reliability Engineer to build observability, reliability, and cost-monitoring for enterprise AI platforms. You will implement OpenTelemetry instrumentation, tracing, SLIs/SLOs, and proactive alerting across AWS-based services and applications.

You’ll collaborate with AI platform, cloud infra, and app teams to ensure reliability, operational visibility, and optimized token usage and compute costs. Strong scripting and IaC skills are essential.

Qualifications

  • 7+ years of experience in platform engineering, SRE, DevOps, or production infrastructure operations.
  • Hands-on experience with OpenTelemetry SDKs, Collector configuration, instrumentation, and telemetry pipelines.
  • Strong experience implementing distributed tracing, metrics, logging, dashboards, and alerting.
  • Experience defining and implementing reliability metrics, SLOs, error budgets, and availability monitoring.

Responsibilities

  • Observability & Instrumentation: Design and implement OpenTelemetry-based instrumentation across AI agents, services, APIs, and platform infrastructure.
  • AI Agent Monitoring: Implement end-to-end tracing for agent execution, LLM interactions, tool calls, multi-agent workflows, and external API integrations.
  • Platform Reliability: Define and implement Service Level Indicators (SLIs), SLOs, availability monitoring, and reliability dashboards across AI platform components.
  • Alerting & Incident Management: Configure proactive alerts for service degradation, execution failures, latency spikes, and resource utilization.
  • Cost Observability: Build monitoring and reporting capabilities for LLM token consumption, model API costs, AWS compute utilization, and agent execution costs.
  • Observability Integration: Integrate telemetry from AI platforms, AWS services, and applications into centralized observability solutions.
  • Automation & Continuous Improvement: Automate observability configuration, dashboard provisioning, and alert management using IaC.

Skills

OpenTelemetry
Distributed tracing
Telemetry pipelines
Python scripting
Terraform / IaC
Cost monitoring
AWS observability

Education

7+ years Platform Engineering / SRE / DevOps

Tools

Prometheus
Grafana
AWS CloudWatch
EKS
OpenTelemetry SDKs

Job description

Newpage Solutions is seeking an experienced Platform Reliability Engineer to build observability, reliability, and cost-monitoring for enterprise AI platforms. You will implement OpenTelemetry instrumentation, tracing, SLIs/SLOs, and proactive alerting across AWS-based services and applications.

You’ll collaborate with AI platform, cloud infra, and app teams to ensure reliability, operational visibility, and optimized token usage and compute costs. Strong scripting and IaC skills are essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Reliability Engineer (Observability & AI Platform)
Platform Reliability Engineer (Observability & AI Platform)

Newpage Solutions • Polska

Remote
PLN 469,000 - 665,000
People-first culture
Smart collaboration
Work-life balance
+2
Senior AI Observability Engineer
Senior AI Observability Engineer

Tealium Inc. • Polska

Hybrid
PLN 290,000 - 375,000
Remote-first working
New-hire equity grants
Stipends for home office
+1
Senior AI SRE - Reliability, Observability & Cost Control
Senior AI SRE - Reliability, Observability & Cost Control

EPAM Systems • Warszawa

Hybrid
PLN 170,000 - 320,000
Health insurance
Multisport
Shopping vouchers
+1
Senior Platform SRE: Reliability & Observability Architect
Senior Platform SRE: Reliability & Observability Architect

IG KnowHow • Kraków

Hybrid
PLN 240,000 - 420,000
Growth opportunities
Mentoring programs
Networking clubs
+2
Remote Observability Engineer - Scale & Reliability
Remote Observability Engineer - Scale & Reliability

Whatnot • Kraków

Hybrid
PLN 520,000 - 580,000
AI-Driven Site Reliability Engineer for Cloud Observability
AI-Driven Site Reliability Engineer for Cloud Observability

Luxoft Poland • Poland

On-site
PLN 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies • Województwo małopolskie

On-site
PLN 133,920 - 200,880
Remote Observability Platform Engineer for Scalable Systems
Remote Observability Platform Engineer for Scalable Systems

Whatnot Inc. • Kraków

Hybrid
PLN 240,000 - 360,000
Senior Observability Engineer - Remote, Scale Reliability
Senior Observability Engineer - Remote, Scale Reliability

Whatnot • Kraków

On-site
PLN 561,000 - 786,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Luxoft • Poland

On-site
PLN 180,000 - 260,000
Private Medical & Dental care
Life Insurance covered
Internal Mobility program