Site Reliability Engineer

Socket.dev

Cambridge (MA)

On-site

USD 76,000 - 136,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Healthcare
401K
Paid time off
Parental leave
Employee assistance program

Job summary

Akamai is hiring a Site Reliability Engineer to improve reliability, performance, and scalability of Compute platforms. You will fix complex Linux/network issues, build automation, and apply AI-assisted tooling to speed incident response.

Ideal candidates have a CS/CE degree, experience with distributed systems, and proficiency in Python or Go, plus hands-on with Prometheus/Grafana/Loki and Docker/Kubernetes or Nomad. This is a US-based role with competitive compensation and benefits.

Qualifications

  • Bachelor's degree in Computer Engineering or Computer Science or equivalent
  • Experience supporting large-scale distributed systems
  • Linux and networking knowledge including routing, DNS, firewalls, TCP/IP, and L7 traffic management
  • Proficient in Python or Go
  • Experience with observability tools such as Prometheus, Grafana, Loki, ELK/OpenSearch, or similar
  • Experience with infrastructure automation or configuration management tools such as Terraform, Ansible, Salt, or similar
  • Familiar with Docker or Podman and orchestration platforms such as Kubernetes or Nomad

Responsibilities

  • Troubleshooting complex issues across Linux systems, networking, and distributed services.
  • Building software and automation to reduce toil and improve efficiency.
  • Developing and applying AI-assisted tooling to accelerate incident investigation and reliability.
  • Using data analysis and network diagnostics to identify improvements.
  • Establishing and improving monitoring, alerting, SLIs, and SLOs for critical services.
  • Contributing to root cause analysis and post-incident reviews.
  • Partnering with Engineering to improve system design and operational readiness.
  • Participating in on-call rotation and leading incident response.

Skills

Linux
Networking
Python or Go
Observability tools
Infra automation
Containers
Kubernetes/Nomad

Education

Bachelor's degree in Computer Engineering or Computer Science or equivalent

Tools

Docker
Podman
Terraform
Ansible
Salt
Kubernetes
Nomad
Prometheus
Grafana
Loki
ELK/OpenSearch

Job description

Are you passionate about cutting edge technology?

Does building next generation Cloud Computing technology excites you?

Join our Compute Site Reliability Engineering team

Our team is responsible for improving the reliability, performance, and scalability of our Compute products and platforms. We solve complex problems, improve how our systems operate, and build automation that makes our platforms more resilient.

Partner with the best

In this role, you'll work at the intersection of systems, networking, and software engineering to solve problems across globally distributed services. You'll investigate complex production behavior, turn operational insights into lasting engineering improvements, and help shape how our services are operated.

As a Site Reliability Engineer, you will be responsible for:

  • Troubleshooting complex issues across Linux systems, networking, and distributed services.
  • Building software and automation that reduce operational toil, improve efficiency, and prevent recurring issues.
  • Developing and applying AI-assisted tooling to accelerate incident investigation, identify operational patterns, reduce toil, and improve reliability.
  • Using data analysis, network diagnostics, and debugging tools to identify performance and reliability improvements.
  • Establishing and improve monitoring, alerting, SLIs, and SLOs for critical services.
  • Contributing to root cause analysis, post-incident reviews, and long-term corrective actions.
  • Partnering with Engineering teams to improve system design, deployment safety, and operational readiness.
  • Participating in an on-call rotation and providing leadership during incident response, driving timely service restoration, effective communication, and post-incident improvement efforts.

Do what you love

To be successful in this role you will:

  • Have relevant experience and a Bachelor's degree in Computer Engineering, Computer Science or equivalent
  • Have experience supporting large-scale distributed systems
  • Have Linux and networking knowledge, including routing, DNS, firewalls, TCP/IP, and L7 traffic management
  • Be proficient in a programming language such as Python or Go
  • Have experience with observability tools such as Prometheus, Grafana, Loki, ELK/OpenSearch, or similar
  • Have experience with infrastructure automation or configuration management tools such as Terraform, Ansible, Salt, or similar.
  • Be familiar with container technologies such as Docker or Podman and orchestration platforms such as Kubernetes or Nomad.

About us

At Akamai, we make life better for billions of people, trillions of times a day. Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.

Our focus is simple: Cloud and Edge: Running apps closer to users for instant performance. Security: Neutralizing threats before they ever reach your data. Content Delivery: Scaling the world's biggest moments without a glitch. AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.

At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.

We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Compensation

Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $75,700 - $136,300/year; a candidate’s salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary.

  • The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP).
  • Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness.

Eligibility requirements apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies GmbH • Cambridge (MA)

On-site
USD 121,000 - 219,000
Healthcare
401K
Paid time off
+3
Site Reliability Engineer
Site Reliability Engineer

Akamai Technologies GmbH • Cambridge (MA)

On-site
USD 75,700 - 136,300
Health insurance
401K savings plan
Parental leave
+1
Senior Manager Engineering (Akamai Cloud Experience)
Senior Manager Engineering (Akamai Cloud Experience)

Akamai Technologies GmbH • Cambridge (MA)

On-site
USD 173,000 - 314,000
Healthcare
401K savings plan
Paid time off (PTO)
+2
Senior Manager Engineering (Akamai Cloud Experience)
Senior Manager Engineering (Akamai Cloud Experience)

Akamai Career Site • United States

On-site
USD 174,000 - 313,000
Healthcare
401K
Paid time off
+3
Senior Manager Engineering (Akamai Cloud Experience)
Senior Manager Engineering (Akamai Cloud Experience)

Akamai Technologies • Cambridge (MA)

Hybrid
USD 173,000 - 314,000
Healthcare
401K
Paid time off
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies • Cambridge (MA)

On-site
USD 146,000 - 264,000
Health benefits
401K savings plan
Paid time off (PTO)
Senior Technical Enablement Architect
Senior Technical Enablement Architect

Akamai Technologies • Cambridge (MA)

On-site
USD 91,000 - 164,000
FlexBase program
Healthcare benefits
PTO
Site Reliability Engineer II
Site Reliability Engineer II

Akamai Technologies • Cambridge (MA)

Remote
USD 95,000 - 171,000
Flexible working options
Healthcare benefits
401K savings plan
+2
Site Reliability Engineer
Site Reliability Engineer

Akamai Technologies • Cambridge (MA)

Hybrid
USD 75,000 - 137,000
Healthcare
401(k) plan
PTO
+1
Site Reliability Engineer II
Site Reliability Engineer II

Akamai Technologies GmbH • Cambridge (MA), Northern (KY)

On-site
USD 95,000 - 171,000
Flexible working options
Health insurance
401K savings plan