Site Reliability Engineer (SRE) – GCP Platform

ITC Infotech

Bengaluru

On-site

INR 900,000 - 1,300,000

Full time

11 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ITC Infotech is seeking a Site Reliability Engineer (SRE) for the GCP Platform in Bangalore. The role requires 8+ years of hands-on experience and a strong focus on reliability, on-call rotations, and incident management.

You will design, implement and operate cloud infrastructure on GCP with Kubernetes, Terraform, and observability tooling. You will collaborate with cross-functional teams to improve system resilience and drive automation across CI/CD pipelines, monitoring, and incident response

Qualifications

  • Experience in designing and operating large-scale cloud platforms.
  • Strong background in SRE practices, incident response and on-call.
  • Hands-on with IaC, CI/CD, and observability tools.

Responsibilities

  • Own end-to-end production reliability with a focus on availability, performance and cost.
  • Lead incident response, post-incident reviews and runbook improvements.
  • Design and operate infrastructure on GCP, including GKE, IAM and networking.
  • Implement and manage observability, monitoring and alerting.
  • Collaborate with engineering teams to improve system resilience.

Skills

GCP
Kubernetes
Python
CI/CD
Linux

Education

Bachelor's degree in CS/Eng

Tools

Terraform
GKE
Dynatrace/Grafana

Job description

Role - Site Reliability Engineer (SRE) – GCP Platform

Location - Bangalore (O Shaughnessy Road)

Work type - Work from Office

Experience - 8+ Years

Key Responsibilities

Reliability & Operations

  • Own end-to-end production systems reliability, availability, scalability, cost and performance.
  • Drive measurable improvements in MTTR, MTTA, and incident response practices using automation and runbook additions and process enhancements.
  • Participate in 24x7 on-call rotations and handle high-severity incidents and document the learnings on ongoing basis.
  • Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics for mission critical services and partner with engineering teams with full accountability for upholding the SLOs.
  • Partner with the various engineering, operations and cloud management teams to deliver highly reliable service in a timely manner.
  • Design, deploy, and manage infrastructure on Google Cloud Platform (GCP).
  • Work extensively on:
  • Compute, networking, IAM, Load Balancers, TLS Certs
  • BigQuery, Pub/Sub, cloud logging enhancement, metrics and logs analysis
  • Implement and manage infrastructure using Terraform (Infrastructure as Code).
  • Deploy and manage containerized workloads using Kubernetes (GKE).
  • Troubleshoot issues related to:
  • Pods, nodes, networking, storage, services on an ongoing basis
  • Manage deployments using Helm, YAML, and rollout strategies (Canary/Blue-Green).

Automation & CI/CD

  • Build and maintain CI/CD pipelines using:
  • Jenkins (pipeline-based, Groovy / Shell / Python scripting)
  • Strong experience in using GitHub as a PowerUser
  • Develop automation using Python and Shell scripting.
  • Reduce operational toil through automation initiatives.

Observability & Monitoring

  • Implement and manage monitoring systems using:
  • Dynatrace, Grafana, logs and metrics explorer
  • Work with logs, metrics, and traces for deep observability to identify trends and arrest problems proactively.
  • Define alerting strategies based on system behaviour and SLOs and create runbooks.
  • Work alongside operations teams to identify, fix the production incidents and own the problem resolution.
  • Work with engineering teams to isolate infra and application issues and set up right tooling for debugging production incidents.

System & Application Troubleshooting

  • Distributed systems
  • Microservices-based architectures on containerised workloads
  • Java and Golang applications
  • Application issues
  • Infrastructure issues
  • Network-related problems

Plan and execute continuous improvement

  • Identify and eliminate repetitive manual tasks.
  • Drive reliability engineering practices and culture. (DRY – Don’t Repeat Yourself)
  • Collaborate with development teams to improve system design and resilience.

Technical Skills

  • Strong expertise in Google Cloud Platform (GCP):
  • GKE, VPC, IAM, Load Balancing, LB, Certs, KMS, logs and metrics exploration
  • BigQuery, Pub/Sub
  • Good understanding of cloud architecture and landing zones

Infrastructure as Code

  • Strong hands-on experience with Terraform
  • Ability to write and debug Terraform code from scratch

Containers & Orchestration

  • Deep expertise in:
  • Docker
  • Strong troubleshooting experience in Kubernetes environments

CI/CD & Automation

  • Hands-on experience with:
  • Jenkins (pipeline-based CI/CD)
  • GitHub
  • Python (preferred)
  • Shell scripting
  • Experience with automation frameworks and tooling

Observability

  • Experience with:
  • Dynatrace / Grafana
  • Log, metrics, and trace-based monitoring
  • Working knowledge of:
  • Java and/or Golang applications
  • Strong debugging skills across application and infrastructure layers
  • Deep understanding of TCP/IP networking
  • Ability to debug network issues in distributed systems

Reliability Engineering Skills

  • Solid understanding of:
  • SLI, SLO, SLA, Error Budgets
  • Demonstrable and Proven Experience improving:
  • MTTR, MTTA
  • Experience handling incident management lifecycle

Soft Skills

  • Strong analytical and troubleshooting mindset
  • Excellent communication and stakeholder management
  • Ability to work in high-pressure production environments
  • Ownership-driven and proactive approach

Preferred candidates with:

  • Exposure to banking/financial domain (optional but valuable)
  • Understanding of security and compliance practices
  • Experience with deployment strategies:
  • Canary, Blue-Green
  • Hands‑on production troubleshooting expert
  • Good at automation + reducing toil
  • Deep understanding of SRE principles
  • Comfortable in 24x7 production environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering
Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering

Deloitte & Touche GmbH Wirtschaftsprüfungsgesellschaft • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Site Reliability Engineer (SRE) - Google Cloud Platform
Site Reliability Engineer (SRE) - Google Cloud Platform

Aziro • Hyderabad

Hybrid
INR 1,500,000 - 3,200,000
Intermediate Applications Developer
Intermediate Applications Developer

UPS • Chennai District

On-site
INR 1,500,000 - 2,000,000
Site Reliability Engineering (SRE) -GCP Devops-Hyderabad.Chennai,Bangalore-6 to 10 yrs
Site Reliability Engineering (SRE) -GCP Devops-Hyderabad.Chennai,Bangalore-6 to 10 yrs

Tata Consultancy Services • Bengaluru

On-site
INR 800,000 - 1,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MangoApps • Maharashtra

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

UST • Pune District

On-site
INR 1,800,000 - 3,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MangoApps INC. • Pune District

On-site
INR 1,400,000 - 1,800,000
SRE
SRE

Jobtailor • Chennai District

On-site
INR 2,500,000 - 4,500,000
Site Reliability Engineer
Site Reliability Engineer

Indihire Consultants • Hyderabad

Hybrid
INR 1,500,000 - 2,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities