Senior Site Reliability Engineer

Akamai Technologies GmbH

Cambridge (MA)

On-site

USD 121,000 - 219,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Healthcare
401K
Paid time off
Sick time
Parental leave
Employee assistance program

Job summary

Akamai is seeking an AI Hardware Site Reliability Engineer to oversee next-generation AI hardware infrastructure, ensuring uptime and reliability across global private cloud and on-prem environments.

You will develop tooling in Python, implement IaC, build telemetry and dashboards, and participate in 24x7 on-call rotations to keep systems responsive and available.

Qualifications

  • Experience developing tooling in Python for automation.
  • Familiarity with infrastructure-as-code practices.
  • Experience with monitoring and incident response.

Responsibilities

  • Develop and scale programmatic tooling and IaC utilities in Python to automate provisioning.
  • Integrate automated workflows across ticketing systems to improve time-to-mitigate for hardware events.
  • Leverage AI utilities and LLM-assisted development to speed scripting and analysis.
  • Work on private cloud and compute tech to improve availability and latency of high-density hardware.
  • Design telemetry pipelines and Prometheus/Grafana dashboards with anomaly detection.
  • Participate in 24x7 on-call rotations and manage incident response via PagerDuty and Slack.
  • Collaborate with third-party vendors and on-site technicians to maintain uptime.

Skills

Python
Infrastructure as Code
Telemetry
Incident management

Tools

Prometheus
Grafana

Job description

Do you enjoy collaborating with teams to solve complex challenges?
Do you enjoy solving large scale distributed content delivery challenges?
Join our critical AI Hardware SRE Team!

The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings.

Partner with the best

In this role, you'll play a part in pioneering the reliability an elite, high-density hardware and software infrastructure spanning the globe. You'll collaborate with product teams from the earliest stages of development to ensure the reliability, scalability, and performance of our systems. You'll define key performance indicators and defend them when they are breached.

Do what you love

To be successful in this role you will:

  • Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning.
  • Integrating automated workflows across disconnected corporate ticketing systems to optimize time-to-mitigate metrics for hardware and network break-fix events.
  • Leveraging advanced AI utilities and LLM-assisted development paradigms where appropriate to accelerate technical execution, script authorship, and system analysis
  • Working on cutting-edge private cloud and compute technologies to improve the availability, latency, and overall systemic health of high-density hardware environments.
  • Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection tailored for bare-metal and virtualized environments.
  • Participating 24x7x365 on-call rotations, spearheading real-time incident management, and managing high-severity service disruption protocols via automated PagerDuty and Slack workflows.
  • Partnering directly with third-party infrastructure vendors and coordinating on-site field technicians to facilitate uptime activities.
About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai:

We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Compensation

Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $121,400 - $218,600/year; a candidate’s salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies, Inc. • Little Rock (AR)

On-site
USD 121,400 - 218,600
Healthcare
401K savings plan
PTO
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies, Inc. • Phoenix (AZ)

On-site
USD 121,400 - 218,600
Healthcare
401K savings plan
Paid time off
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies, Inc. • Hartford (CT)

On-site
USD 121,400 - 218,600
Healthcare
401K savings plan
Paid time off
+2
Site Reliability Engineer
Site Reliability Engineer

Akamai Technologies, Inc. • Santa Fe (NM)

Hybrid
USD 75,700 - 136,300
Healthcare
401K savings plan
Vacation (PTO)
+2
Site Reliability Engineer
Site Reliability Engineer

Akamai Technologies, Inc. • Richmond (VA)

Hybrid
USD 75,700 - 136,300
Healthcare
401K savings plan
Parental leave
Site Reliability Engineer
Site Reliability Engineer

Akamai Technologies, Inc. • Hartford (CT)

Hybrid
USD 75,700 - 136,300
Comprehensive healthcare benefits
401K savings plan
Generous PTO policy
+1
Site Reliability Engineer
Site Reliability Engineer

Akamai Technologies, Inc. • Phoenix (AZ)

On-site
USD 75,700 - 136,300
Healthcare
401K savings plan
Parental leave
+1
Site Reliability Engineer
Site Reliability Engineer

Akamai Technologies, Inc. • Atlanta (GA)

Hybrid
USD 75,700 - 136,300
Healthcare
401K savings plan
Parental leave
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Socket.dev • Cambridge (MA)

Hybrid
USD 146,000 - 264,000
Site Reliability Engineer
Site Reliability Engineer

Akamai Technologies GmbH • Cambridge (MA)

Hybrid
USD 75,000 - 137,000
Health insurance
401K savings plan
Parental leave
+1