Site Reliability Engineer - Google Cloud Platform

GAMMASTACK

Kolkata District

On-site

INR 1,800,000 - 2,800,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

GAMMASTACK is seeking a seasoned Site Reliability Engineer to own the reliability, availability, and performance of production workloads on Google Cloud Platform. You will design resilient infrastructure using Cloud Run and Kubernetes, implement strong observability, and lead incident response to keep services healthy.

The role emphasizes CI/CD improvements, automation, and collaboration with product teams to ensure scalable, secure, and reliable systems.

Qualifications

  • 5+ years of experience in Site Reliability Engineering or related DevOps with production ownership.
  • Strong experience running production systems on Google Cloud Platform.
  • Hands-on with Cloud Run, Kubernetes, and container-based microservices in production.
  • Experience with infrastructure as code (Terraform/ Terragrunt).
  • Strong observability understanding (OpenTelemetry, Cloud Monitoring, New Relic).
  • Experience building or improving CI/CD pipelines and release workflows (GitHub Actions).
  • Ability to code/automate in Python or Java.
  • Excellent written and verbal communication across teams.
  • Experience with AI tooling and agentic workflows is a plus.
  • Experience in retail/e-commerce is a plus.

Responsibilities

  • Own reliability, availability, and performance of microservices and production workloads.
  • Design and improve resilient infrastructure on GCP, with strong emphasis on Cloud Run, Kubernetes, and containerized services.
  • Build observability across logs, metrics, tracing, alerting, and service health so issues are detected early and resolved quickly.
  • Improve deployment safety through stronger CI/CD pipelines, release controls, rollback strategies, and environment consistency.
  • Lead incident response and production readiness practices, including runbooks, postmortems, on-call hygiene, capacity planning, and resilience testing.
  • Reduce operational toil by automating repetitive work and improving tooling for engineers supporting distributed services.
  • Partner with development teams to improve the operability, scalability, and fault tolerance of microservices early in the design lifecycle.
  • Strengthen cloud security and infrastructure hygiene across IAM, secrets management, workload hardening, and production safeguards.
  • Improve service performance, resource efficiency, and cloud cost management without compromising reliability.
  • Support architecture and reliability reviews for critical services and high-traffic business.

Skills

Site Reliability Engineering
DevOps
Google Cloud Platform
Cloud Run
Kubernetes
Terraform
Terragrunt
Observability
CI/CD
GitHub Actions
Python
Java
Incident response

Tools

Terraform
Terragrunt
OpenTelemetry
New Relic
Cloud Monitoring

Job description

  • Own the reliability, availability, and performance of microservices and production workloads.
  • Design and improve resilient infrastructure on GCP, with strong emphasis on Cloud Run, Kubernetes, and containerized services.
  • Build and maintain observability across logs, metrics, tracing, alerting, and service health so issues are detected early and resolved quickly.
  • Improve deployment safety through stronger CI/CD pipelines, release controls, rollback strategies, and environment consistency.
  • Lead incident response and production readiness practices, including runbooks, postmortems, on-call hygiene, capacity planning, and resilience testing.
  • Reduce operational toil by automating repetitive work and improving tooling for engineers supporting distributed services.
  • Partner with development teams to improve the operability, scalability, and fault tolerance of microservices early in the design lifecycle.
  • Strengthen cloud security and infrastructure hygiene across IAM, secrets management, workload hardening, and production safeguards.
  • Improve service performance, resource efficiency, and cloud cost management without compromising reliability.
  • Support architecture and reliability reviews for critical services and high-traffic business 5+ years of experience in Site Reliability Engineering or closely related DevOps roles with meaningful production ownership.
  • Strong experience running production systems on Google Cloud Platform.
  • Hands-on experience with Cloud Run, Kubernetes, and container-based microservices in production.
  • Strong experience with infrastructure as code, particularly Terraform and Terragrunt.
  • Strong understanding of observability using tools such as OpenTelemetry, Cloud Monitoring, New Relic, or equivalent systems.
  • Strong understanding of distributed systems, microservice failure modes, reliability engineering, and production debugging.
  • Experience building or improving CI/CD pipelines and release workflows in modern engineering environments, including GitHub Actions.
  • Ability to write code and automation in one or more languages such as Python or Java.
  • Good judgment during incidents and a practical mindset around reliability, recovery, and risk tradeoffs.
  • Strong written and verbal communication skills, with the ability to work effectively across engineering teams.
  • Experience working with AI tooling and agentic workflows in engineering or operational environments.
  • Experience in retail, e-commerce, or other customer-facing environments is a plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

MangoApps INC. • Pune District

On-site
INR 1,400,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MangoApps • Maharashtra

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer (SRE) – GCP
Senior Site Reliability Engineer (SRE) – GCP

Opstree Global • Bengaluru

On-site
INR 3,600,000 - 6,000,000
Lead DevOps Engineer
Lead DevOps Engineer

V2 Solutions • Khordha

On-site
INR 3,000,000 - 5,400,000
Gammastack - DevOps Engineer - CI/CD Pipeline
Gammastack - DevOps Engineer - CI/CD Pipeline

GAMMASTACK • Kolkata District

On-site
INR 1,200,000 - 2,000,000
Site Reliability Engineer (SRE) – GCP Platform
Site Reliability Engineer (SRE) – GCP Platform

ITC Infotech • Bengaluru

On-site
INR 900,000 - 1,300,000
Site Reliability Engineer- GCP
Site Reliability Engineer- GCP

Aziro • Hyderabad

Hybrid
INR 1,400,000 - 2,100,000
Site Reliability Engineer
Site Reliability Engineer

Advance Career Solutions • Pune District

On-site
INR 1,200,000 - 1,800,000
Site Reliability Engineer (SRE) - Google Cloud Platform
Site Reliability Engineer (SRE) - Google Cloud Platform

Aziro • Hyderabad

Hybrid
INR 1,500,000 - 3,200,000
Associate Senior Platform Engineer
Associate Senior Platform Engineer

Quantiphi Analytics Solutions • Bangalore Rural, Bengaluru, Mumbai

Hybrid
INR 1,200,000 - 2,200,000