Hybrid Site Reliability Engineer — AI Tools & Cloud Edge

Socket.dev

Cambridge (MA)

On-site

USD 76,000 - 136,000

Full time

10 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Healthcare
401K
Paid time off
Parental leave
Employee assistance program

Job summary

Akamai is hiring a Site Reliability Engineer to improve reliability, performance, and scalability of Compute platforms. You will fix complex Linux/network issues, build automation, and apply AI-assisted tooling to speed incident response.

Ideal candidates have a CS/CE degree, experience with distributed systems, and proficiency in Python or Go, plus hands-on with Prometheus/Grafana/Loki and Docker/Kubernetes or Nomad. This is a US-based role with competitive compensation and benefits.

Qualifications

  • Bachelor's degree in Computer Engineering or Computer Science or equivalent
  • Experience supporting large-scale distributed systems
  • Linux and networking knowledge including routing, DNS, firewalls, TCP/IP, and L7 traffic management
  • Proficient in Python or Go
  • Experience with observability tools such as Prometheus, Grafana, Loki, ELK/OpenSearch, or similar
  • Experience with infrastructure automation or configuration management tools such as Terraform, Ansible, Salt, or similar
  • Familiar with Docker or Podman and orchestration platforms such as Kubernetes or Nomad

Responsibilities

  • Troubleshooting complex issues across Linux systems, networking, and distributed services.
  • Building software and automation to reduce toil and improve efficiency.
  • Developing and applying AI-assisted tooling to accelerate incident investigation and reliability.
  • Using data analysis and network diagnostics to identify improvements.
  • Establishing and improving monitoring, alerting, SLIs, and SLOs for critical services.
  • Contributing to root cause analysis and post-incident reviews.
  • Partnering with Engineering to improve system design and operational readiness.
  • Participating in on-call rotation and leading incident response.

Skills

Linux
Networking
Python or Go
Observability tools
Infra automation
Containers
Kubernetes/Nomad

Education

Bachelor's degree in Computer Engineering or Computer Science or equivalent

Tools

Docker
Podman
Terraform
Ansible
Salt
Kubernetes
Nomad
Prometheus
Grafana
Loki
ELK/OpenSearch

Job description

Akamai is hiring a Site Reliability Engineer to improve reliability, performance, and scalability of Compute platforms. You will fix complex Linux/network issues, build automation, and apply AI-assisted tooling to speed incident response.

Ideal candidates have a CS/CE degree, experience with distributed systems, and proficiency in Python or Go, plus hands-on with Prometheus/Grafana/Loki and Docker/Kubernetes or Nomad. This is a US-based role with competitive compensation and benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — Scalable Cloud Infra & Automation
Site Reliability Engineer — Scalable Cloud Infra & Automation

Akamai Technologies, Inc. • Honolulu (HI)

On-site
USD 76,000 - 136,000
Healthcare
401(k) plan
Paid time off
+3
Site Reliability Engineer
Site Reliability Engineer

Socket.dev • Cambridge (MA)

On-site
USD 76,000 - 136,000
Healthcare
401K
Paid time off
+2
Site Reliability Engineer: AI-Powered Cloud Platform
Site Reliability Engineer: AI-Powered Cloud Platform

Akamai • United States

Hybrid
USD 120,000 - 180,000
FlexBase program
Site Reliability Engineer II
Site Reliability Engineer II

Akamai Technologies GmbH • Cambridge (MA), Northern (KY)

On-site
USD 95,000 - 171,000
Flexible working options
Health insurance
401K savings plan
Remote Senior Site Reliability Engineer - Cloud & Edge
Remote Senior Site Reliability Engineer - Cloud & Edge

Akamai Career Site • United States

On-site
USD 147,000 - 264,000
Healthcare
401K
Paid time off
+3
Site Reliability Engineer: Cloud Automation & CI/CD
Site Reliability Engineer: Cloud Automation & CI/CD

Akamai Career Site • United States

On-site
USD 76,000 - 136,000
Remote SRE Engineer - Scale, Automate, Equity Eligible
Remote SRE Engineer - Scale, Automate, Equity Eligible

Akamai Technologies • Cambridge (MA)

Hybrid
USD 75,000 - 137,000
Healthcare
401(k) plan
PTO
+1
Site Reliability Engineer (Guardicore AI Platform) - Remote
Site Reliability Engineer (Guardicore AI Platform) - Remote

Akamai • United States

Hybrid
USD 120,000 - 180,000
FlexBase program
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Remote SRE: Automation & Scalable Infrastructure
Remote SRE: Automation & Scalable Infrastructure

Akamai Technologies GmbH • Cambridge (MA)

Hybrid
USD 75,700 - 136,300
Health insurance
401K savings plan
Parental leave
+1