Senior Site Reliability Engineer

Akamai Technologies

Kraków

On-site

PLN 180,000 - 270,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Akamai Technologies in Kraków, Poland, is looking for a Senior SRE to oversee scalable AI hardware infrastructure, ensuring high availability and reliability of our services.

The role focuses on building tooling in Python, automating provisioning, and integrating workflows across Jira, Siebel, and PagerDuty, while advancing observability and incident response capabilities.

Qualifications

  • Formal CS education or equivalent practical experience in large‑scale SRE or Production Engineering roles.
  • Experience building automation tools and API integrations.
  • Hands-on with observability stacks and timeseries engines such as Prometheus, Grafana, OpenTelemetry, Loki.

Responsibilities

  • Developing and scaling robust programmatic tooling and infrastructure‑as‑code utilities in Python to eliminate operational toil and automate fleet‑wide provisioning.
  • Integrating automated workflows across ticketing platforms like JIRA, Siebel, and PagerDuty to enhance resolution times for hardware and network issues.
  • Leveraging AI utilities and LLM‑assisted development to improve scripting, execution, and system evaluation.
  • Enhancing private cloud and compute technologies to optimize availability, latency, and health in high‑density hardware environments.
  • Designing telemetry pipelines, Prometheus/Grafana dashboards, and AI‑based anomaly detection for bare‑metal and virtualized environments.
  • Participating in 24x7x365 on‑call rotations, incident management, and blameless post‑mortems with automated PagerDuty and Slack workflows.
  • Partnering with third‑party infrastructure vendors and coordinating on‑site field technicians to facilitate uptime activities.

Skills

Python
SRE
Automation

Education

Bachelor's degree in CS

Tools

Prometheus
Grafana
OpenTelemetry
Loki

Job description

Job Description

Join our critical AI Hardware SRE Team! The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best‑in‑class uptime and reliability of our AI hardware infrastructure offerings.

Responsibilities
  • Developing and scaling robust programmatic tooling and infrastructure‑as‑code utilities in Python to eliminate operational toil and automate fleet‑wide provisioning.
  • Integrating automated workflows across disparate ticketing platforms like JIRA, Siebel, and PagerDuty to enhance resolution times for hardware and network issues.
  • Leveraging advanced AI utilities and LLM‑assisted development paradigms to enhance technical execution, script creation, and system evaluation effectively.
  • Enhancing advanced private cloud and compute technologies to consistently optimize availability, latency, and systemic health in high‑density hardware environments.
  • Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI‑based anomaly detection tailored for bare‑metal and virtualized environments.
  • Participating in 24x7x365 on‑call rotations, spearheading real‑time incident management, and managing high‑severity service disruption protocols via automated PagerDuty and Slack workflows.
  • Partnering directly with third‑party infrastructure vendors and coordinating on‑site field technicians to facilitate uptime activities.
Qualifications
  • Possess a solid Computer Science foundation, demonstrated through formal education or equivalent practical experience in large‑scale SRE or Production Engineering roles.
  • Demonstrate proficient tooling and coding ability in languages like Python to build scalable operational tools, API integrations, and automation frameworks.
  • Show hands‑on experience with modern observability stacks and timeseries engines, like Prometheus, Grafana, OpenTelemetry, and Loki.
  • Possess a working understanding of advanced networking topologies, high‑bandwidth routing/switching infrastructure, BGP, and dual‑stack IPv4/IPv6 networks.
  • Demonstrate expertise designing service rollouts, establishing operational readiness criteria, telemetry baselines, and defining effective alerting thresholds.
  • Demonstrate extensive experience building technical runbooks, leading complex incident response bridges, and driving comprehensive, blameless post‑mortems.
  • Demonstrate a proven ability to fully own ambiguous technical challenges, coordinate cross‑functional teams, and diligently pursue production‑grade solutions.
Benefits

Benefits at Akamai: We support your health, well‑being, finances, and life beyond work. See our benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies • Województwo małopolskie

On-site
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies GmbH • Kraków

Hybrid
PLN 90,000 - 130,000
Flexible working options
Health benefits
Professional development opportunities
Senior II Site Reliability Engineer
Senior II Site Reliability Engineer

Akamai Technologies • Kraków

On-site
PLN 300,000 - 520,000
Benefits at Akamai
FlexBase program
Hybrid/work flexibility
Senior II Site Reliability Engineer
Senior II Site Reliability Engineer

Akamai Technologies GmbH • Kraków

On-site
PLN 200,000 - 320,000
FlexBase program
Comprehensive benefits
Senior Site Reliability Engineer (Core Linux Platforms) - Remote
Senior Site Reliability Engineer (Core Linux Platforms) - Remote

Akamai Technologies • Kraków

On-site
PLN 180,000 - 260,000
FlexBase work flexibility
Health and wellbeing benefits
Senior Site Reliability Engineer (Cloud and Networking) - Remote
Senior Site Reliability Engineer (Cloud and Networking) - Remote

Akamai Technologies • Kraków

On-site
PLN 90,000 - 130,000
Health benefits
Financial support
Family support
+1
Senior Site Reliability Engineer (Server Enablement & Qualification) - Remote
Senior Site Reliability Engineer (Server Enablement & Qualification) - Remote

Akamai Technologies • Kraków

Hybrid
PLN 80,000 - 110,000
Comprehensive benefits
Flexibility to work from home or in-office
Hybrid work model
Senior Systems Site Reliability Engineer, B2B
Senior Systems Site Reliability Engineer, B2B

Jobtailor • Poland

On-site
PLN 180,000 - 320,000
Senior Site Reliability Engineer (Linux Performance) - Remote
Senior Site Reliability Engineer (Linux Performance) - Remote

Akamai Technologies • Kraków

Hybrid
PLN 120,000 - 160,000
Health benefits
Well-being support
FlexBase program
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Grid Dynamics • Województwo pomorskie

On-site
PLN 80,000 - 120,000
Medical insurance
Sports benefits
Professional development opportunities
+2