Site Reliability Engineer (GCP)

Persistent

Hyderabad

On-site

INR 2,500,000 - 5,000,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health check-ups
Hybrid work option
Office accessibility

Job summary

Persistent is seeking a Site Reliability Engineer (GCP) to enhance reliability, scalability, and security of cloud-native services in Hyderabad. You will design and operate enterprise CI/CD platforms, manage production workloads on GKE/Cloud Run, and drive automation to improve MTTR and reliability.

The role requires 8–12 years of SRE/DevOps experience with strong GCP expertise, IaC, and robust incident management skills.

Qualifications

  • 8–12 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Operations, or related disciplines.
  • Strong hands-on experience building CI/CD pipelines with Jenkins, GitHub Actions, GitLab CI, Argo CD, Tekton, or Cloud Build.
  • Extensive hands-on experience with Google Cloud Platform (GCP) including GKE and Cloud Run.
  • Strong expertise with Google Kubernetes Engine (GKE) and Kubernetes-based deployments.
  • Experience managing Cloud Run, Cloud Build, IAM, VPC Networking, Monitoring, and Logging services.
  • Experience deploying and operating containerized applications using Docker and Kubernetes.
  • Proficiency in scripting/automation using Python, Bash, Go, or similar languages.
  • Strong understanding of Git workflows, YAML, source code management, and release management.
  • Experience supporting mission-critical production environments and incident response.
  • Expertise in observability, monitoring, logging, tracing, and cloud-native practices.
  • Experience defining SLIs/SLOs/Error Budgets and reliability metrics.
  • Experience with Infrastructure as Code using Terraform or similar.
  • Understanding of Policy-as-Code and infrastructure compliance standards.
  • Experience with OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, Splunk.
  • Hands-on with automation frameworks, self-healing, ChatOps, and runbooks.
  • Strong troubleshooting and root cause analysis skills across distributed systems.
  • Knowledge of capacity planning, performance optimization, and resilience testing.
  • Understanding cloud security principles and operational best practices.
  • Ability to collaborate with developers, architects, platform teams, and stakeholders.
  • Excellent communication and documentation skills.
  • GCP certifications or equivalent experience is a plus.
  • Passion for reliability engineering and operational excellence.
  • Bachelor's degree in Computer Science, IT, Engineering, or equivalent.

Responsibilities

  • Design, build, and maintain secure, scalable CI/CD pipelines across Development, Test, and Production.
  • Monitor, troubleshoot, and resolve pipeline failures, deployment issues, and environment drift.
  • Implement deployment automation, progressive delivery, rollback mechanisms, and IaC practices.
  • Build, deploy, and support production workloads on GKE and Cloud Run.
  • Manage containerized environments focusing on scalability, availability, networking, and security.
  • Coordinate deployments with development teams and manage release readiness.
  • Provide production support and lead incident management including triage and post-incident reviews.
  • Define and enhance observability through metrics, logs, traces, dashboards, and SLIs/SLOs.
  • Perform root cause analysis and implement preventive reliability improvements.
  • Build POCs and automation to reduce toil and improve MTTR.
  • Apply SRE practices including capacity planning, resilience testing, and operational readiness.

Skills

CI/CD pipelines
GCP
Kubernetes
Incident management
Automation
Observability
Security
SRE practices
Terraform
Python/Bash/Go

Education

Bachelor's degree in CS/IT/Engineering

Tools

Jenkins
GitHub Actions
GitLab CI
Argo CD
Tekton
Google Cloud Build
GKE
Cloud Run
Terraform
OpenTelemetry

Job description

About Position:

We are seeking a highly skilled Site Reliability Engineer (SRE) with strong expertise in Google Cloud Platform (GCP) to enhance the reliability, scalability, security, and operational excellence of cloud-native services. The ideal candidate will be responsible for building and operating enterprise-scale CI/CD platforms, managing production workloads on Google Kubernetes Engine (GKE) and Cloud Run, implementing observability solutions, and driving automation initiatives that improve reliability, performance, and operational efficiency. This role requires deep expertise in cloud-native technologies, infrastructure automation, incident management, and modern SRE practices.



  • Role: Site Reliability Engineer (GCP)

  • Location: Hyderabad

  • Experience: 8 to 12 Years

  • Job Type: Full-Time Employment


What You'll Do:


  • Design, build, and maintain secure, scalable, and reliable CI/CD pipelines across Development, Test, and Production environments.

  • Monitor, troubleshoot, and resolve pipeline failures, deployment issues, release bottlenecks, and environment drift.

  • Implement deployment automation, progressive delivery strategies, rollback mechanisms, release gates, and Infrastructure-as-Code practices.

  • Build, deploy, and support production workloads on Google Kubernetes Engine (GCP) and Cloud Run.

  • Manage containerized application environments with a focus on scalability, availability, networking, configuration management, and security.

  • Coordinate application deployments with development teams and manage release readiness activities.

  • Provide production support and lead incident management activities including triage, escalation, communication, recovery, and post-incident reviews.

  • Define and enhance observability through metrics, logs, traces, dashboards, alerting, SLIs, SLOs, and Error Budgets.

  • Perform root cause analysis and implement preventive measures and long-term reliability improvements.

  • Build proof‑of‑concepts and automation solutions that reduce operational toil and improve Mean Time to Recovery (MTTR).

  • Apply Site Reliability Engineering practices including capacity planning, resilience testing, reliability reviews, and operational readiness assessments.

  • Collaborate with globally distributed engineering, platform, and operations teams.

  • Drive continuous improvement initiatives focused on automation, reliability, scalability, and operational excellence.

  • Maintain platform security, governance, compliance, and cloud infrastructure best practices.


Expertise You'll Bring:


  • 8 to 12 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Operations, or related disciplines.

  • Strong hands‑on experience building and maintaining CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, Argo CD, Tekton, and Google Cloud Build.

  • Extensive hands‑on experience with Google Cloud Platform (GCP).

  • Strong expertise with Google Kubernetes Engine (GKE) and Kubernetes‑based application deployments.

  • Experience managing Cloud Run, Cloud Build, Artifact Registry, IAM, VPC Networking, Cloud Monitoring, and Cloud Logging services.

  • Strong experience deploying and operating containerized applications using Docker and Kubernetes.

  • Proficiency in scripting and automation using Python, Bash, Go, or similar programming languages.

  • Strong understanding of Git workflows, YAML, source code management, and release management practices.

  • Experience supporting mission‑critical production environments and managing incident response processes.

  • Expertise in observability, monitoring, logging, tracing, and cloud‑native operational practices.

  • Hands‑on experience defining and managing Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and reliability metrics.

  • Experience implementing Infrastructure as Code using Terraform or similar frameworks.

  • Understanding of Policy‑as‑Code, governance controls, and infrastructure compliance standards.

  • Experience with observability platforms such as OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, and Splunk.

  • Hands‑on experience implementing automation frameworks, self‑healing solutions, ChatOps workflows, and operational runbooks.

  • Strong troubleshooting and root cause analysis skills across distributed and cloud‑native systems.

  • Experience with capacity planning, performance optimization, scalability engineering, and resilience testing.

  • Strong understanding of cloud security principles and operational best practices.

  • Ability to work effectively with developers, architects, platform teams, business stakeholders, and distributed delivery teams.

  • Excellent communication, documentation, and stakeholder management skills.

  • Google Cloud certifications or equivalent enterprise GCP experience will be an added advantage.

  • Passion for reliability engineering, cloud automation, continuous improvement, and operational excellence.

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.


Benefits:


  • Competitive salary and benefits package

  • Culture focused on talent development with quarterly growth opportunities and company‑sponsored higher education and certifications

  • Opportunity to work with cutting‑edge technologies

  • Employee engagement initiatives such as project parties, flexible work hours, and Long Service awards

  • Annual health check‑ups

  • Insurance coverage: group term life, personal accident, and Mediclaim hospitalization for self, spouse, two children, and parents


Values‑Driven, People‑Centric & Inclusive Work Environment:

Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds.



  • We support hybrid work and flexible hours to fit diverse lifestyles.

  • Our office is accessibility‑friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities.

  • If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment


"Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind."

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (GCP)
Site Reliability Engineer (GCP)

Persistent Systems Limited • Hyderabad

Hybrid
INR 4,200,000 - 6,600,000
Hybrid work
Flexible hours
Education sponsorship
+3
Devops GCP Engineer
Devops GCP Engineer

Persistent • Hyderabad

Hybrid
INR 1,400,000 - 2,300,000
Hybrid work model
Professional development opportunities
Health check-ups and insurance
Programmer (Dev)-DevOps Lead
Programmer (Dev)-DevOps Lead

Persistent Systems • Hyderabad

On-site
INR 2,500,000 - 4,500,000
Hybrid work model
Flexible hours
Group term life insurance
+2
Devops GCP Engineer
Devops GCP Engineer

Persistent Systems • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Group term life
Personal accident
Mediclaim hospitalization
+3
Programmer (Dev)-DevOps Lead
Programmer (Dev)-DevOps Lead

Persistent • Hyderabad

On-site
INR 4,000,000 - 6,000,000
Competitive salary
Education sponsorship
Cutting-edge technologies
+2
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Persistent Systems Limited • Pune District

On-site
INR 1,400,000 - 2,200,000
Hybrid work
Long Service awards
Company-sponsored education
GCP DevOps Engineer
GCP DevOps Engineer

Persistent • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Competitive salary
Talent development
Cutting-edge technologies
+3
Support Engineer
Support Engineer

Persistent Systems • Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Hybrid work model
Flexible work hours
Long service awards
+4
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Persistent Systems Limited • Pune District

On-site
INR 2,500,000 - 4,200,000
Competitive salary
Benefits package
Talent development
+4
Infrastructure Architect
Infrastructure Architect

TymblHub • Pune District

On-site
INR 4,500,000 - 7,000,000
Competitive salary
Hybrid work options
Healthcare benefits