Senior Site Reliability Engineer

Akamai Career Site

India

On-site

INR 900,000 - 1,400,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

FlexBase program
Hybrid work options

Job summary

Akamai is seeking a Senior Site Reliability Engineer to join its AI hardware SRE team in India. You will build scalable tooling, drive automation, and own proactive reliability for high-density hardware in regional data centers.

You will design telemetry, implement AI-based anomaly detection, and lead incident bridges with cross-functional teams. A 5+ year background and a CS-related degree are required.

Qualifications

  • 5+ years of relevant experience in SRE/DevOps.
  • Bachelor's degree in Computer Science or related field.
  • Proficient in Python for building scalable tooling.
  • Experience with observability stacks and timeseries databases.

Responsibilities

  • Develop programmatic tooling and IaC to automate fleet provisioning.
  • Integrate automated workflows across ticketing systems to speed up incident resolution.
  • Leverage AI tools and LLM-based approaches to improve script generation and system evaluation.
  • Work on private cloud and compute tech to improve availability and latency.
  • Design telemetry pipelines and dashboards with Prometheus/Grafana and AI anomaly detection.
  • Participate in 24x7 on-call rotations and incident response processes.
  • Coordinate with third-party vendors and on-site technicians to maintain uptime.

Skills

Python
Automation tooling
Observability stacks
APIs
AI tooling

Education

Bachelor's degree in Computer Science or related field

Tools

Prometheus
Grafana
OpenTelemetry
Loki

Job description

Do you enjoy collaborating with teams to solve complex challenges?

Do you enjoy solving large scale distributed content delivery challenges?

Join our critical AI Hardware SRE Team!

The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings.

Partner with the best

This position focuses on enhancing system reliability, scalability, and performance across high-density hardware and software infrastructure in regional data centers. Responsibilities include defining KPIs, proactive monitoring, automation, and urgent issue resolution. Collaboration with teams ensures best practices, reduced downtime, and optimized systems. The role supports seamless operations, business-critical applications, and improved user experiences through efficient, data-driven solutions.

As a Senior Site Reliability Engineer, you will be responsible for:

  • Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning.
  • Integrating automated workflows across diverse corporate ticketing systems to enhance resolution times for hardware and network break-fix incidents.
  • Leveraging advanced AI tools and LLM-based development approaches to enhance technical execution, script creation, and comprehensive system evaluation.
  • Working on cutting-edge private cloud and compute technologies to improve the availability, latency, and overall systemic health of high-density hardware environments.
  • Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection tailored for bare-metal and virtualized environments.
  • Participating in 24x7x365 on-call rotations, spearheading real-time incident management, and managing high-severity service disruption protocols via automated PagerDuty and Slack workflows.
  • Partnering directly with third-party infrastructure vendors and coordinating on-site field technicians to facilitate uptime activities

Do what you love

To be successful in this role you will:

  • Have 5+ years of relevant experience and a Bachelor's degree in Computer Science or related field
  • Demonstrate exceptional proficiency in tooling and coding using languages like Python to build scalable operational tools, API integrations, and automation frameworks.
  • Show hands-on experience with modern observability stacks and timeseries engines, like Prometheus, Grafana, OpenTelemetry, and Loki.
  • Possess a working understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks.
  • Demonstrate expertise as a primary designer for new service rollouts, establishing operational readiness criteria, telemetry baselines, and alerting thresholds.
  • Demonstrate extensive experience building technical runbooks, leading complex incident response bridges, and driving comprehensive, blameless post-mortems.
  • Demonstrate a proven ability to fully own ambiguous technical challenges, coordinate cross-functional teams, and drive toward production-grade solutions effectively.

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager Engineering
Senior Manager Engineering

Akamai Technologies • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Software Engineer
Software Engineer

Akamai Technologies GmbH • Bengaluru

Hybrid
INR 700,000 - 1,100,000
Senior Manager Engineering
Senior Manager Engineering

Akamai Technologies GmbH • India

On-site
INR 4,000,000 - 7,000,000
Health benefits
FlexBase program
Career development opportunities
Network Infrastructure Engineer
Network Infrastructure Engineer

Akamai Technologies GmbH • Bengaluru

Hybrid
INR 800,000 - 1,400,000
FlexBase program
Health benefits
Senior Software Engineer (DevOps)
Senior Software Engineer (DevOps)

Akamai Technologies • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Health benefits
Flexible working options
Family and financial support
Principal Software Engineer Lead
Principal Software Engineer Lead

Akamai Technologies • Bengaluru

On-site
INR 4,000,000 - 9,000,000
FlexBase program
Hybrid/remote-friendly work model
Senior Security Engineer
Senior Security Engineer

Akamai Technologies GmbH • India

Hybrid
INR 4,000,000 - 7,000,000
Senior Major Account Executive, Existing Accounts
Senior Major Account Executive, Existing Accounts

Akamai Technologies GmbH • India

Hybrid
INR 6,000,000 - 9,000,000
Senior Major Account Executive, Existing Accounts
Senior Major Account Executive, Existing Accounts

Akamai • India

Hybrid
INR 3,500,000 - 5,200,000
FlexBase program
Hybrid work model
Senior Security Engineer
Senior Security Engineer

Akamai Technologies • Bengaluru

Hybrid
INR 3,000,000 - 5,000,000
FlexBase program