Forward Deployment Engineer (SRE)

PwC Acceleration Centers

Hyderabad

On-site

INR 2,500,000 - 4,000,000

Full time

8 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

PwC Acceleration Centers is seeking a senior, customer‑facing Site Reliability Engineer who owns the reliability of client production systems end to end. Embedded as the main technical contact, you will design reliable infrastructure, lead incidents, apply AI‑assisted operations, and mentor the team.

You will manage SLIs/SLOs, drive blameless post‑mortems, and lead in on‑call rotations, while advancing Kubernetes and AWS operations with AI SRE guardrails.

Qualifications

  • Substantial SRE experience with ownership of production reliability on AWS.
  • Experience running Kubernetes and AWS services at scale.
  • Defining and managing SLIs, SLOs, and error budgets.
  • Strong observability practice.
  • Incident leadership and post‑mortem facilitation.
  • Automation skills in Python, Go, or Bash.
  • Hands‑on AI SRE tooling with guardrails.

Responsibilities

  • Own the reliability of production systems for enterprise customers on AWS.
  • Define and manage SLIs, SLOs, and error budgets to guide reliability.
  • Lead response to P0/P1 incidents and conduct blameless post‑mortems.
  • Lead within on‑call rotation and incident triage.
  • Operate and optimize Kubernetes clusters and AWS services.
  • Introduce AI SRE tooling to speed triage and resolution with guardrails.
  • Advise customers on reliability practices and mentor engineers.
  • Share field learnings with product and engineering teams.

Skills

Observability
Incident leadership
Automation
Python
Go
Bash
AI SRE tooling
AWS
Kubernetes
Terraform
GitOps

Education

AWS Associate-level certification

Tools

Prometheus
Grafana
Datadog
OpenTelemetry
Loki
Tempo
Jaeger
ELK
PagerDuty
Opsgenie
Terraform modules
GitHub Actions
GitLab CI
ArgoCD

Job description

Site Reliability Engineering | Forward Deployed Engineering

Experience Required

6–9 years.

Job Summary

A senior, customer-facing SRE who owns the reliability of a client's production systems end to end. Embedded as the main technical point of contact, you will design reliable infrastructure, lead incidents, apply AI-assisted operations, and mentor the team.

Key Responsibilities
  • Own the reliability of production systems for one or more enterprise customers on AWS.
  • Define and manage SLIs, SLOs, and error budgets, and ensure monitoring provides meaningful insight.
  • Lead the response to serious (P0/P1) incidents and run blameless post-mortems that lead to fixes.
  • Lead within the on‑call rotation.
  • Operate and optimise Kubernetes clusters and AWS services as workloads grow.
  • Introduce AI SRE tooling and AIOps where it speeds up triage and resolution, with appropriate guardrails.
  • Advise customers on improving their reliability practices, and mentor associate engineers.
  • Share field learnings with product and engineering teams.
Required Qualifications
  • Substantial SRE experience with real ownership of production reliability on AWS.
  • Experience running Kubernetes and AWS services at scale.
  • A track record of defining and managing SLIs, SLOs, and error budgets.
  • Strong observability practice.
  • Proven incident leadership and post‑mortem facilitation.
  • Strong automation skills (Python, Go, or Bash).
  • Hands‑on understanding of AI‑assisted operations, including introducing AI SRE tooling with sensible guardrails.
  • AWS certification at Associate level as a minimum.
Preferred Qualifications
  • Owning CI/CD pipelines and release automation.
  • Writing and maintaining Terraform modules.
  • GitOps workflows and Helm.
  • Designing telemetry collection across services.
  • Incident‑management tooling such as PagerDuty.
  • Hands‑on experience integrating AIOps or AI SRE tooling into production operations.
  • Familiarity with observability or evaluation of AI systems (e.g., Langfuse).
  • Experience delivering AI‑based work or AI/ML‑driven initiatives in production.
  • AWS Certified Solutions Architect – Professional and/or AWS Certified DevOps Engineer – Professional.
Technical Skills & Tools
  • Observability & monitoring: Prometheus, Grafana, Datadog, OpenTelemetry, Loki, Tempo/Jaeger (tracing), ELK/Elastic
  • Reliability practices: SLIs/SLOs, error budgets, incident command, capacity planning, failure‑mode analysis
  • Incident & on‑call: PagerDuty, Opsgenie, blameless post‑mortems
  • Automation & scripting: Python, Go, Bash, Git
  • AI‑assisted operations: AI SRE tooling, AIOps, anomaly detection, automated RCA
  • DevOps (good to have): CI/CD (GitHub Actions, GitLab CI), Terraform modules, GitOps (ArgoCD), Helm
  • Mentors and raises standards across the team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC • Hyderabad, Bengaluru

Hybrid
INR 900,000 - 1,400,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Site Reliability Engineer
Site Reliability Engineer

Innodata Inc. • India

On-site
INR 2,400,000 - 4,000,000
Lead SRE
Lead SRE

Cvent, Inc. • India

On-site
INR 2,500,000 - 4,500,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Lead SRE
Lead SRE

Baazi Games • New Delhi

On-site
INR 3,500,000 - 6,000,000
Site Reliability Engineer
Site Reliability Engineer

Cybage Software • Pune District

On-site
INR 2,500,000 - 4,000,000
SRE Lead
SRE Lead

Acldigital • Ahmedabad District

On-site
INR 1,500,000 - 2,000,000
AWS SRE Professional
AWS SRE Professional

Infosys • Bengaluru

On-site
INR 900,000 - 1,300,000