Staff Site Reliability Engineer

PowerToFly

Gurgaon

On-site

INR 1,800,000 - 3,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

PowerToFly seeks a seasoned Site Reliability Engineer/DevOps professional to own and maintain highly available production systems, lead incident response, conduct RCA/PIRs, and drive reliability, performance, and operational excellence.

You will design, build, and manage scalable AWS infrastructure with Terraform, Kubernetes (EKS), and strong ownership of networking, security, and platform resilience; develop CI/CD and GitOps pipelines with GitLab CI and ArgoCD; and mentor teammates while

Qualifications

  • 5–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles.
  • Strong hands-on experience with AWS, Terraform, Kubernetes, and ArgoCD in production environments.
  • Experience operating EKS-based platforms, including networking, scaling, monitoring, and troubleshooting.
  • Strong knowledge of CI/CD, GitOps, and automation practices, with hands-on use of GitLab CI and ArgoCD.
  • Experience managing production systems in a 24×7 environment, including incident response and on-call practices.
  • Solid Linux and cloud networking background.
  • Experience with observability tools such as ELK and Prometheus.
  • Strong scripting skills in Python, Bash, or Go.

Responsibilities

  • Own and maintain production systems, lead incident response (P1/P2), conduct RCA/PIRs, and drive reliability improvements.
  • Design, build, and manage scalable AWS infrastructure using Terraform and Kubernetes (EKS), focusing on network, security, and resilience.
  • Develop and optimize CI/CD and GitOps pipelines using GitLab CI and ArgoCD, while automating operational processes.
  • Manage observability and on-call operations with tools like PagerDuty/Zenduty, Prometheus, Grafana, ELK, and Datadog.
  • Collaborate with global teams and mentor members, contributing to cloud architecture and compliance initiatives and creating operational docs.

Skills

SRE/DevOps mindset
Incident response
Automation
Linux proficiency
Cloud networking

Education

Engineering degree in computer science or equivalent

Tools

AWS
Terraform
Kubernetes (EKS)
ArgoCD
GitLab CI
CI/CD
GitOps
Prometheus
Grafana
ELK
Datadog
Python
Bash
Go

Job description

What You Will Do
  • Own and maintain highly available production systems, lead incident response (P1/P2), conduct RCA/PIRs, and drive improvements to reliability, performance, and operational excellence.
  • Design, build, and manage scalable cloud infrastructure on AWS using Terraform, with strong ownership of Kubernetes (EKS), networking, security, and platform resilience.
  • Develop and optimize CI/CD and GitOps pipelines using GitLab CI and ArgoCD, while automating operational processes to improve efficiency and consistency.
  • Manage observability and on‑call operations through tools such as PagerDuty/Zenduty, Prometheus, Grafana, ELK, and Datadog, ensuring actionable monitoring and effective alert management.
  • Collaborate with global engineering, security, and product teams, contribute to cloud architecture and compliance initiatives (SOC2, ISO27001), create operational documentation, and mentor team members.
What You Will Need
  • 5–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles.
  • Strong hands‑on experience with AWS, Terraform, Kubernetes, and ArgoCD in production environments.
  • Experience operating EKS‑based platforms, including networking, scaling, monitoring, and troubleshooting.
  • Strong knowledge of CI/CD, GitOps, and automation practices, with hands‑on use of GitLab CI and ArgoCD.
  • Experience managing production systems in a 24×7 environment, including incident response and on‑call practices.
  • Solid Linux and cloud networking background.
  • Experience with observability tools such as ELK and Prometheus.
  • Strong scripting skills in Python, Bash, or Go.
Preferred Qualifications
  • Engineering degree in computer science or equivalent.
  • Cloud certifications such as AWS SysOps/DevOps Engineer, or CKA/CKAD.
  • Exposure to ITSM or change management processes in regulated industries (healthcare, fintech, or similar).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Stryker Group • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Saama • Chennai District

On-site
INR 1,200,000 - 1,800,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iLink Digital • Chennai

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Five9 • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Associate Senior Site Reliability Engineer
Associate Senior Site Reliability Engineer

Global Payments Inc. • Pune District

On-site
INR 1,500,000 - 2,100,000
Site Reliability Engineer
Site Reliability Engineer

Innodata Inc. • India

On-site
INR 2,400,000 - 4,000,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Arcesium • Hyderabad, Bengaluru

Hybrid
INR 6,000,000 - 9,000,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Headout • Bengaluru

On-site
INR 1,200,000 - 1,800,000