Senior Site Reliability Engineer

Akamai Career Site

Poland

Hybrid

PLN 180,000 - 320,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

FlexBase program

Job summary

Akamai Krakow is seeking a Senior Site Reliability Engineer to help scale and maintain our AI hardware and distributed systems. You will design programmatic tooling, automate provisioning, and drive reliability across regional data centers.

You will collaborate with product teams, implement telemetry pipelines, and maintain Prometheus/Grafana dashboards. Strong Python, automation, and networking knowledge are essential for on-call incident response in a fast-moving AI infrastructure.

Qualifications

  • Solid CS foundation or equivalent practical experience in large-scale SRE/Production Engineering.
  • Proficient in Python to build scalable tooling and automations.
  • Hands-on experience with modern observability stacks (Prometheus, Grafana, OpenTelemetry, Loki).
  • Understanding of high-bandwidth networking topologies and IPv4/IPv6.59

Responsibilities

  • Develop and scale programmatic tooling and IaC utilities in Python.
  • Integrate automated workflows across ticketing platforms (JIRA, Siebel, PagerDuty).
  • Leverage AI utilities and LLM-assisted development to improve execution and automation.
  • Enhance private cloud and compute tech to optimize availability and latency.
  • Design telemetry pipelines and dashboards (Prometheus/Grafana) for bare-metal and virtual environments.
  • Participate in 24x7 on-call rotations and manage incident response and post-mortems.
  • Coordinate with third-party vendors and on-site field technicians to maintain uptime.

Skills

Python
Observability
Incident management
Networking (BGP/IPv6)
Cross-functional teamwork
On-call readiness

Education

Computer Science degree

Tools

Prometheus
Grafana
OpenTelemetry
Loki
JIRA
PagerDuty
Siebel
Slack

Job description

Do you enjoy collaborating with teams to solve complex challenges?

Do you enjoy solving large scale distributed content delivery challenges?

Join our critical AI Hardware SRE Team!

The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings.

Partner with the best

This position involves ensuring reliability and operational readiness for advanced hardware and software systems across regional data centers. Collaborate with product teams during development to optimize scalability, performance, and system reliability. Define and monitor key performance indicators while addressing breaches. Analytical abilities, coding expertise, urgent issue resolution, and dedication to maintaining service availability are essential for success in this role.

As a Senior Site Reliability Engineer, you will be responsible for:

  • Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning.
  • Integrating automated workflows across disparate ticketing platforms like JIRA, Siebel, and PagerDuty to enhance resolution times for hardware and network issues.
  • Leveraging advanced AI utilities and LLM-assisted development paradigms to enhance technical execution, script creation, and system evaluation effectively.
  • Enhancing advanced private cloud and compute technologies to consistently optimize availability, latency, and systemic health in high-density hardware environments.
  • Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection tailored for bare-metal and virtualized environments.
  • Participating in 24x7x365 on-call rotations, spearheading real-time incident management, and managing high-severity service disruption protocols via automated PagerDuty and Slack workflows.
  • Partnering directly with third-party infrastructure vendors and coordinating on-site field technicians to facilitate uptime activities.

Do what you love

To be successful in this role you will:

  • Possess a solid Computer Science foundation, demonstrated through formal education or equivalent practical experience in large-scale SRE or Production Engineering roles.
  • Demonstrate proficient tooling and coding ability in languages like Python to build scalable operational tools, API integrations, and automation frameworks.
  • Show hands-on experience with modern observability stacks and timeseries engines, like Prometheus, Grafana, OpenTelemetry, and Loki.
  • Possess a working understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks.
  • Demonstrate expertise designing service rollouts, establishing operational readiness criteria, telemetry baselines, and defining effective alerting thresholds.
  • Demonstrate extensive experience building technical runbooks, leading complex incident response bridges, and driving comprehensive, blameless post-mortems.
  • Demonstrate a proven ability to fully own ambiguous technical challenges, coordinate cross-functional teams, and diligently pursue production-grade solutions.

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Job Info
  • Locations Centrum Biurowe Vinci, Krakow, Malopolskie, 31-323, PL (Remote)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Remote
Senior Site Reliability Engineer - Remote

Akamai Career Site • Poland

Hybrid
PLN 250,000 - 360,000
FlexBase adaptivity
Home/office hybrid flexibility
Senior II Site Reliability Engineer
Senior II Site Reliability Engineer

Akamai Career Site • Poland

Hybrid
PLN 180,000 - 320,000
Senior II Site Reliability Engineer
Senior II Site Reliability Engineer

Akamai Technologies GmbH • Kraków

On-site
PLN 200,000 - 320,000
FlexBase program
Comprehensive benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies GmbH • Kraków

Hybrid
PLN 90,000 - 130,000
Flexible working options
Health benefits
Professional development opportunities
Senior II Site Reliability Engineer
Senior II Site Reliability Engineer

Akamai Technologies • Kraków

On-site
PLN 300,000 - 520,000
Benefits at Akamai
FlexBase program
Hybrid/work flexibility
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies • Kraków

On-site
PLN 180,000 - 270,000
Software Engineer II (Data Products)
Software Engineer II (Data Products)

Akamai Technologies GmbH • Kraków

Hybrid
PLN 180,000 - 240,000
FlexBase-like workplace program
Benefits for health and wellbeing
Principal Software Engineer
Principal Software Engineer

Akamai Technologies GmbH • Kraków

Hybrid
PLN 127,000 - 213,000
Flexible working options
Comprehensive health benefits
Financial and family support programs
Software Engineer (Frontend)
Software Engineer (Frontend)

Akamai Technologies GmbH • Kraków

Hybrid
PLN 192,000 - 279,000
Health support
Work flexibility
Senior Devops Engineer
Senior Devops Engineer

Akamai Technologies, Inc. • Poland

Remote
PLN 260,000 - 380,000