Site Reliability Engineer II

Akamai Technologies, Inc.

Richmond (VA)

On-site

USD 95,000 - 171,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Healthcare benefits
401K savings plan
Paid Time Off
Parental leave
Employee assistance program

Job summary

Akamai Technologies, Inc. is looking for a Site Reliability Engineer II to join their Inference Cloud Team in Richmond, Virginia. The role focuses on automating and monitoring AI platforms, requiring expertise in Linux systems, Python or Go programming, and familiarity with Kubernetes and CI/CD pipelines.

Candidates should have at least 2 years of experience in Site Reliability Engineering and a strong desire to learn about AI infrastructure. Akamai offers innovative compensation and benefits to support employee health and wellness.

Qualifications

  • 2+ years of experience in Site Reliability Engineering.
  • Demonstrated coding ability in Python or Go.
  • Experience with Linux systems administration.
  • Familiarity with Kubernetes and containerization.
  • Experience with monitoring tools like Prometheus or Grafana.
  • Exposure to CI/CD pipelines and infrastructure-as-code tools.

Responsibilities

  • Maintain dashboards and monitoring for inference workloads.
  • Write automation and tooling in Python or Go.
  • Contribute to SLO tracking and reporting.
  • Support CI/CD pipeline maintenance and deployment safety.
  • Collaborate with product teams for troubleshooting.
  • Participate in on-call rotations for production incidents.

Skills

Site Reliability Engineering
Python
Go
Linux systems administration
Kubernetes
Monitoring tools
CI/CD
Infrastructure as code

Education

Bachelor's Degree in a relevant field (or equivalent experience)

Tools

Prometheus
Grafana
Terraform
SaltStack

Job description

Are you passionate about cutting-edge AI infrastructure?

Do you want to build your SRE career on one of the most exciting platforms in cloud computing?

Join the Akamai Inference Cloud Team

The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications.

In this role, responsibilities will include automation, monitoring, incident response, and working collaboratively with skilled team members. Candidates should possess expertise in Linux systems, automation, and SRE practices. Daily activities involve coding, improving dashboards, enhancing alerts, and minimizing repetitive tasks. Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform.

As an Site Reliability Engineer II, you will be responsible for:
  • Building and maintaining dashboards, alerts, and monitoring for inference workloads using Akamai's existing observability platform
  • Writing automation and tooling in Python or Go to reduce operational toil and improve system reliability
  • Building and improving runbooks for inference-specific operational procedures, integrating into Akamai's existing incident management processes
  • Contributing to SLO tracking and reporting, identifying trends and areas for improvement
  • Supporting CI/CD pipeline maintenance, deployment safety checks, and rollback procedures
  • Collaborating with product engineering teams to troubleshoot complex problems across the stack
  • Participating in on‑call rotations, responding to production incidents, and conducting blameless post‑mortems
To be successful in this role you will:
  • Have 2+ years of experience in Site Reliability Engineering and a Bachelor's Degree or its equivalent experience
  • Demonstrate coding ability in at least one programming language (Python or Go) with experience writing automation
  • Have experience with Linux systems administration and the ability to troubleshoot complex infrastructure issues
  • Show familiarity with Kubernetes and containerization concepts
  • Have experience with monitoring and observability tools such as Prometheus, Grafana, or similar
  • Have exposure to CI/CD pipelines and infrastructure‑as‑code tools (Terraform, SaltStack, or equivalent)
  • Show a willingness to learn and grow, with genuine curiosity about AI infrastructure and distributed systems
Benefits
  • Your health
  • Your finances
  • Your family
  • Your time at work
  • Your time pursuing other endeavors
Compensation

Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $95,000 - $171,000/year; a candidate’s salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry‑leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.

Equal Employment Opportunity Rights

Akamai Technologies is an affirmative action, equal opportunity employer that values the strength that diversity brings to the workplace. All qualified applicants will receive consideration for employment and will not be discriminated against on the basis of gender, gender identity, sexual orientation, race/ethnicity, protected veteran status, disability, or other protected group status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

Akamai Technologies, Inc. • Hartford (CT)

On-site
USD 95,000 - 171,000
Health benefits
401K savings plan
Parental leave
+1
Site Reliability Engineer II
Site Reliability Engineer II

Akamai Technologies, Inc. • Columbus (OH)

On-site
USD 95,000 - 171,000
Healthcare
401K savings plan
Parental leave
+1
Site Reliability Engineer II
Site Reliability Engineer II

Akamai Technologies • Cambridge (MA)

Remote
USD 95,000 - 171,000
Flexible working options
Healthcare benefits
401K savings plan
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies, Inc. • Hartford (CT)

On-site
USD 121,400 - 218,600
Healthcare
401K savings plan
Paid time off
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies, Inc. • Annapolis (MD)

On-site
USD 121,400 - 218,600
Healthcare
401K savings plan
Employee Stock Purchase Plan (ESPP)
+2
Site Reliability Engineer
Site Reliability Engineer

Akamai Technologies, Inc. • Hartford (CT)

Hybrid
USD 75,700 - 136,300
Comprehensive healthcare benefits
401K savings plan
Generous PTO policy
+1
Site Reliability Engineer
Site Reliability Engineer

Akamai Technologies, Inc. • Richmond (VA)

Hybrid
USD 75,700 - 136,300
Healthcare
401K savings plan
Parental leave
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies GmbH • Cambridge (MA)

On-site
USD 121,000 - 219,000
Healthcare
401K
Paid time off
+3
Site Reliability Engineer
Site Reliability Engineer

Akamai Technologies GmbH • Cambridge (MA)

Hybrid
USD 75,000 - 137,000
Health insurance
401K savings plan
Parental leave
+1
Senior II Site Reliability Engineer
Senior II Site Reliability Engineer

Akamai Technologies, Inc. • Salt Lake City (UT)

On-site
USD 146,400 - 263,600