Site Reliability Engineer: AI-Powered Cloud Platform

Akamai

United States

Hybrid

USD 120,000 - 180,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

FlexBase program

Job summary

Akamai is seeking a Site Reliability Engineer to own the reliability and performance of the Guardicore Data and AI Platform. You will operate secure Kubernetes infra for core microservices, data pipelines, and internal tooling, while driving observability, security, and cost efficiency.

You'll lead complex production investigations, mentor engineers, and leverage AI to auto-remediate incidents. Collaboration across DevOps, Software, Data, AI and Security teams is essential, with on-call

Qualifications

  • 3+ years of SRE/DevOps/platform engineering experience with complex systems.
  • Ability to design and implement monitoring and observability strategies.
  • Production experience with Kubernetes, Docker, Helm, and multi-cloud environments.
  • Experience with GitOps, CI/CD, and Infrastructure as Code.
  • Scripting in Python, Go, and Bash; comfortable with AI tooling.

Responsibilities

  • Operate secure, highly available Kubernetes infrastructure for microservices and data pipelines.
  • Enhance platform reliability, observability, security, performance, and cost efficiency.
  • Provide guidance to engineers to improve service reliability.
  • Lead complex production investigations and drive long-term improvements.
  • Leverage AI tools to auto-remediate incidents and streamline operations.
  • Collaborate across DevOps, Software, Data, AI, and Security teams; participate in on-call rotations.

Skills

SRE
DevOps
Platform engineering
Monitoring
Python
Go
Bash
AI tooling

Tools

Prometheus
Grafana
Kubernetes
Docker
Helm
CI/CD
GitOps
IaC

Job description

Akamai is seeking a Site Reliability Engineer to own the reliability and performance of the Guardicore Data and AI Platform. You will operate secure Kubernetes infra for core microservices, data pipelines, and internal tooling, while driving observability, security, and cost efficiency.

You'll lead complex production investigations, mentor engineers, and leverage AI to auto-remediate incidents. Collaboration across DevOps, Software, Data, AI and Security teams is essential, with on-call

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (Guardicore AI Platform) - Remote
Site Reliability Engineer (Guardicore AI Platform) - Remote

Akamai • United States

Hybrid
USD 120,000 - 180,000
FlexBase program
Site Reliability Engineer: Cloud Automation & CI/CD
Site Reliability Engineer: Cloud Automation & CI/CD

Akamai Career Site • United States

On-site
USD 76,000 - 136,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Hybrid Site Reliability Engineer — AI Tools & Cloud Edge
Hybrid Site Reliability Engineer — AI Tools & Cloud Edge

Socket.dev • Cambridge (MA)

On-site
USD 76,000 - 136,000
Healthcare
401K
Paid time off
+2
Site Reliability Engineer — Scalable Cloud Infra & Automation
Site Reliability Engineer — Scalable Cloud Infra & Automation

Akamai Technologies, Inc. • Honolulu (HI)

On-site
USD 76,000 - 136,000
Healthcare
401(k) plan
Paid time off
+3
Site Reliability Engineer
Site Reliability Engineer

Akamai Technologies • Cambridge (MA)

Hybrid
USD 75,000 - 137,000
Healthcare
401(k) plan
PTO
+1
Site Reliability Engineer II
Site Reliability Engineer II

Akamai Technologies GmbH • Cambridge (MA), Northern (KY)

On-site
USD 95,000 - 171,000
Flexible working options
Health insurance
401K savings plan
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Global Network SRE: Scale, Reliability & Automation
Global Network SRE: Scale, Reliability & Automation

Akamai • United States

On-site
USD 146,000 - 264,000
FlexBase program
Healthcare
401K savings plan
+2
Senior SRE: AI Cloud Platform & Kubernetes
Senior SRE: AI Cloud Platform & Kubernetes

Lambda Inc. • San Francisco (CA)

Hybrid
USD 190,000 - 270,000
Health insurance
401k with company match
Flexible PTO
+2