Site Reliability Engineer

PwC Acceleration Center India

Bengaluru

On-site

INR 2,200,000 - 3,800,000

Full time

15 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

PwC Acceleration Center India is seeking a senior Site Reliability Engineer to own the reliability of production systems end-to-end for enterprise clients on AWS. You will be embedded as the main technical contact, designing resilient infrastructure, leading incidents, and mentoring the team with AI-assisted operations.

You will define SLIs/SLOs, manage error budgets, and drive blameless post-mortems while mentoring associates and sharing insights with product teams.

Qualifications

  • Substantial SRE experience with production reliability on AWS.
  • Experience running Kubernetes and AWS services at scale.
  • Defined and managed SLIs, SLOs, and error budgets.
  • Strong observability practice.
  • Incident leadership and post-mortem facilitation.
  • Automation skills (Python, Go, or Bash).
  • Hands-on AI-assisted operations including guardrails.

Responsibilities

  • Own the reliability of production systems for one or more enterprise customers on AWS.
  • Define and manage SLIs, SLOs, and error budgets, ensuring meaningful monitoring.
  • Lead response to serious incidents and blameless post-mortems.
  • Lead within the on-call rotation.
  • Operate and optimise Kubernetes clusters and AWS services as workloads grow.
  • Introduce AI SRE tooling and AIOps for faster triage and resolution with guardrails.
  • Advise customers on reliability practices and mentor engineers.
  • Share field learnings with product and engineering teams.

Skills

AWS
Kubernetes
SLIs/SLOs
Incident leadership
Python
Go
Bash
AI SRE tooling
Terraform

Tools

Prometheus
Grafana
Datadog
OpenTelemetry
Loki
Tempo/Jaeger
ELK/Elastic
PagerDuty
Opsgenie
GitHub Actions
GitLab CI
ArgoCD
Helm
Terraform modules

Job description

Role- Site Reliability Engineering ( Mandatory skills AWS, AI & SLI OR SLO)Certification mandatory
Experience Required -6-8years
Aws or Azure or Devop's Certification mandatory
Job Summary

A senior, customer-facing SRE who owns the reliability of a client's production systems end to end. Embedded as the main technical point of contact, you will design reliable infrastructure, lead incidents, apply AI-assisted operations, and mentor the team.

Key Responsibilities
  • Own the reliability of production systems for one or more enterprise customers on AWS.
  • Define and manage SLIs, SLOs, and error budgets, and ensure monitoring provides meaningful insight.
  • Lead the response to serious (P0/P1) incidents and run blameless post-mortems that lead to fixes.
  • Lead within the on-call rotation.
  • Operate and optimise Kubernetes clusters and AWS services as workloads grow.
  • Introduce AI SRE tooling and AIOps where it speeds up triage and resolution, with appropriate guardrails.
  • Advise customers on improving their reliability practices, and mentor associate engineers.
  • Share field learnings with product and engineering teams.
Required Qualifications
  • Substantial SRE experience with real ownership of production reliability on AWS.
  • Experience running Kubernetes and AWS services at scale.
  • A track record of defining and managing SLIs, SLOs, and error budgets.
  • Strong observability practice.
  • Proven incident leadership and post-mortem facilitation.
  • Strong automation skills (Python, Go, or Bash).
  • Hands‑on understanding of AI-assisted operations, including introducing AI SRE tooling with sensible guardrails.
  • AWS certification at Associate level as a minimum.
Preferred Qualifications
  • Owning CI/CD pipelines and release automation.
  • Writing and maintaining Terraform modules.
  • GitOps workflows and Helm.
  • Designing telemetry collection across services.
  • Incident-management tooling such as PagerDuty.
  • Hands‑on experience integrating AIOps or AI SRE tooling into production operations.
  • Familiarity with observability or evaluation of AI systems (e.g., Langfuse).
  • Experience delivering AI-based work or AI/ML-driven initiatives in production.
  • AWS Certified Solutions Architect – Professional and/or AWS Certified DevOps Engineer – Professional.
Technical Skills & Tools
  • Observability & monitoring: Prometheus, Grafana, Datadog, OpenTelemetry, Loki, Tempo/Jaeger (tracing), ELK/Elastic
  • Reliability practices: SLIs/SLOs, error budgets, incident command, capacity planning, failure‑mode analysis
  • Incident & on‑call: PagerDuty, Opsgenie, blameless post‑mortems
  • Automation & scripting: Python, Go, Bash, Git
  • AI‑assisted operations: AI SRE tooling, AIOps, anomaly detection, automated RCA
  • DevOps (good to have): CI/CD (GitHub Actions, GitLab CI), Terraform modules, GitOps (ArgoCD), Helm
  • Mentors and raises standards across the team.
  • Sound technical judgement.
Experience Required

6–9 years.

Reporting & Team

Embedded with an enterprise customer as their primary technical point of contact, within our Forward Deployed Engineering practice; works alongside client engineers a

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer_ AWS certified
Site Reliability Engineer_ AWS certified

PwC Acceleration Center India • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Software Engineer
Software Engineer

PwC • Hyderabad, Bengaluru

Hybrid
INR 2,800,000 - 5,200,000
Forward Deployed Engineer - Observability & Platform
Forward Deployed Engineer - Observability & Platform

PwC • Hyderabad, Bengaluru

On-site
INR 2,500,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Larsen & Toubro Infotech Ltd (LTI) • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Site Reliability Engineer
Site Reliability Engineer

Hackajob • Ahmedabad District, Gurugram District, Mumbai

On-site
INR 1,800,000 - 2,600,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Solutions By Text • Bengaluru

On-site
INR 800,000 - 1,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities