Site Reliability Engineer_ AWS certified

PwC Acceleration Center India

Bengaluru

On-site

INR 4,000,000 - 6,000,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

PwC Acceleration Center India is seeking a senior Site Reliability Engineer to own the reliability of a client's production systems end to end on AWS. Embedded as the main technical point of contact, you will design reliable infrastructure, lead incidents, and mentor the team.

The role emphasizes defining SLIs and SLOs, AI-assisted operations, and blameless post-mortems to improve stability. You will lead the on-call rotation and guide customers on reliability while collaborating with product

Qualifications

  • Substantial SRE experience with ownership of production reliability on AWS.
  • Experience running Kubernetes and AWS services at scale.
  • Strong observability and incident leadership experience.
  • Hands-on automation skills with Python, Go, or Bash.
  • Experience introducing AI SRE tooling and guardrails.

Responsibilities

  • Own the reliability of production systems for one or more enterprise customers on AWS.
  • Define and manage SLIs, SLOs, and error budgets, and ensure monitoring provides meaningful insight.
  • Lead the response to serious (P0/P1) incidents and run blameless post-mortems that lead to fixes.
  • Lead within the on-call rotation.
  • Operate and optimise Kubernetes clusters and AWS services as workloads grow.
  • Introduce AI SRE tooling and AIOps where it speeds up triage and resolution, with appropriate guardrails.
  • Advise customers on improving their reliability practices, and mentor associate engineers.
  • Share field learnings with product and engineering teams.

Skills

SRE ownership
Incident leadership
Python/Go/Bash
AWS experience
Kubernetes experience
Observability

Tools

Prometheus
Grafana
Datadog
OpenTelemetry
PagerDuty
ArgoCD
Terraform

Job description

Role- Site Reliability Engineering ( Mandatory skills AWS, AI & SLI OR SLO)

Experience Required -6-8years

Job Summary

A senior, customer-facing SRE who owns the reliability of a client's production systems end to end. Embedded as the main technical point of contact, you will design reliable infrastructure, lead incidents, apply AI-assisted operations, and mentor the team.

Key Responsibilities
  • Own the reliability of production systems for one or more enterprise customers on AWS.
  • Define and manage SLIs, SLOs, and error budgets, and ensure monitoring provides meaningful insight.
  • Lead the response to serious (P0/P1) incidents and run blameless post-mortems that lead to fixes.
  • Lead within the on-call rotation.
  • Operate and optimise Kubernetes clusters and AWS services as workloads grow.
  • Introduce AI SRE tooling and AIOps where it speeds up triage and resolution, with appropriate guardrails.
  • Advise customers on improving their reliability practices, and mentor associate engineers.
  • Share field learnings with product and engineering teams.
Required Qualifications
  • Substantial SRE experience with real ownership of production reliability on AWS.
  • Experience running Kubernetes and AWS services at scale.
  • A track record of defining and managing SLIs, SLOs, and error budgets.
  • Strong observability practice.
  • Proven incident leadership and post-mortem facilitation.
  • Strong automation skills (Python, Go, or Bash).
  • Hands-on understanding of AI-assisted operations, including introducing AI SRE tooling with sensible guardrails.
  • AWS certification at Associate level as a minimum.
Preferred Qualifications
  • Owning CI/CD pipelines and release automation.
  • Writing and maintaining Terraform modules.
  • GitOps workflows and Helm.
  • Designing telemetry collection across services.
  • Incident-management tooling such as PagerDuty.
  • Hands-on experience integrating AIOps or AI SRE tooling into production operations.
  • Familiarity with observability or evaluation of AI systems (e.g., Langfuse).
  • Experience delivering AI-based work or AI/ML-driven initiatives in production.
  • AWS Certified Solutions Architect Professional and/or AWS Certified DevOps Engineer – Professional.
Technical Skills & Tools
  • Observability & monitoring: Prometheus, Grafana, Datadog, OpenTelemetry, Loki, Tempo/Jaeger (tracing), ELK/Elastic
  • Reliability practices: SLIs/SLOs, error budgets, incident command, capacity planning, failure-mode analysis
  • Incident & on-call: PagerDuty, Opsgenie, blameless post-mortems
  • Automation & scripting: Python, Go, Bash, Git
  • AI-assisted operations: AI SRE tooling, AIOps, anomaly detection, automated RCA
  • DevOps (good to have): CI/CD (GitHub Actions, GitLab CI), Terraform modules, GitOps (ArgoCD), Helm
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

PwC Acceleration Center India • Bengaluru

On-site
INR 2,200,000 - 3,800,000
Software Engineer
Software Engineer

PwC • Hyderabad, Bengaluru

Hybrid
INR 2,800,000 - 5,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Larsen & Toubro Infotech Ltd (LTI) • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Hackajob • Ahmedabad District, Gurugram District, Mumbai

On-site
INR 1,800,000 - 2,600,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Site Reliability Engineer
Site Reliability Engineer

Solutions By Text • Bengaluru

On-site
INR 800,000 - 1,200,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC • Hyderabad, Bengaluru

On-site
INR 900,000 - 1,400,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Forward Deployed Engineer - Observability & Platform
Forward Deployed Engineer - Observability & Platform

PwC • Hyderabad, Bengaluru

On-site
INR 2,500,000 - 4,200,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000