Senior Site Reliability Engineer

Akamai Technologies

Województwo małopolskie

On-site

PLN 133,920 - 200,880

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

A technology company is seeking a Site Reliability Engineer in Poland to tackle complex reliability challenges in AI workloads. In this role, you'll build observability, improve deployment safety, and partner with teams to enhance reliability across platforms. Required skills include proficiency in Kubernetes, Python, and experience in managing distributed systems. You will also automate workflows and participate in on-call rotations. Join the team if you are passionate about technology and want to contribute to cutting-edge AI solutions.

Qualifications

  • Expertise in SRE, managing large-scale distributed systems.
  • Familiarity with Kubernetes and containerization systems.
  • Proficiency in Python or Go for automation.

Responsibilities

  • Build and maintain observability for AI workloads.
  • Write automation and tooling to reduce operational toil.
  • Collaborate with product teams to improve reliability.

Skills

SRE expertise
Kubernetes
Python
Go
Prometheus
Grafana
Terraform
AI/ML infrastructure

Job description

Overview

Do you enjoy solving complex reliability challenges for cutting-edge technology? Do you have a passion for automation and building systems that scale?

Join the Akamai Inference Cloud Team. The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications with unmatched performance, compliance, and economics.

Responsibilities
  • As a Site Reliability Engineer, you will be responsible for building and maintaining observability for AI workloads, including telemetry, dashboards, alerts, SLO/SLI tracking, and driving improvements when targets are missed.
  • Write automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response.
  • Integrate AI workloads into Akamai's existing incident management processes, build runbooks, participate in on-call rotations, and conduct blameless post-mortems.
  • Build and maintain CI/CD integrations, deployment safety checks, and rollback automation.
  • Collaborate with product engineering teams to improve reliability, contribute to architecture decisions, and ensure operational readiness for product releases.
  • Contribute to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure.
What You Will Love / Qualifications
  • Demonstrate expertise in SRE, infrastructure, or platform engineering, managing large-scale distributed systems with extensive operational experience.
  • Demonstrate expertise in Kubernetes and large-scale containerization systems.
  • Define SLOs and work with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring.
  • Demonstrate proficiency in Python or Go for automation, CI/CD pipelines, deployment safety, and infrastructure-as-code like Terraform.
  • Interest in or experience with AI/ML infrastructure, model serving, or GPU workloads.
  • Resolve issues independently while maintaining accountability throughout the process.
  • Demonstrate accountability for reliability, develop automation and monitoring, and collaborate effectively with an engineering team unfamiliar with SRE practices.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies GmbH • Kraków

On-site
PLN 90,000 - 130,000
Flexible working options
Health benefits
Professional development opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies • Kraków

On-site
PLN 180,000 - 270,000
Senior Site Reliability Engineer - Remote
Senior Site Reliability Engineer - Remote

Akamai Career Site • Poland

Hybrid
PLN 250,000 - 360,000
FlexBase adaptivity
Home/office hybrid flexibility
Senior II Site Reliability Engineer
Senior II Site Reliability Engineer

Akamai Technologies GmbH • Kraków

On-site
PLN 200,000 - 320,000
FlexBase program
Comprehensive benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Career Site • Poland

Hybrid
PLN 180,000 - 320,000
FlexBase program
Senior II Site Reliability Engineer
Senior II Site Reliability Engineer

Akamai Technologies • Kraków

On-site
PLN 300,000 - 520,000
Benefits at Akamai
FlexBase program
Hybrid/work flexibility
Senior II Site Reliability Engineer
Senior II Site Reliability Engineer

Akamai Career Site • Poland

Hybrid
PLN 180,000 - 320,000
Lead AI Platform Engineer
Lead AI Platform Engineer

Akamai Technologies • Kraków

Remote
PLN 299,000 - 386,000
Remote Senior AI Hardware SRE - Automation & Reliability
Remote Senior AI Hardware SRE - Automation & Reliability

Akamai Career Site • Poland

Hybrid
PLN 180,000 - 320,000
FlexBase program
Senior SRE - Compute Cloud Reliability & Automation
Senior SRE - Compute Cloud Reliability & Automation

Akamai Career Site • Poland

Hybrid
PLN 180,000 - 320,000