Site Reliability Engineer

Recro

Bengaluru

On-site

INR 3,500,000 - 6,500,000

Full time

10 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Recro is seeking an SRE to ensure reliability, scalability, availability, and observability of high-traffic production systems. You will automate infrastructure, manage Kubernetes, and build robust CI/CD pipelines.

Ideal candidates have 3+ years in SRE/DevOps, extensive AWS or GCP experience, and strong scripting skills for automation. The role involves on-call shifts and capacity planning, with focus on cost optimization and performance.

Qualifications

  • 3+ years in SRE/DevOps or Infra engineering in production
  • Experience in high-traffic or large-scale environments
  • Strong AWS or GCP experience
  • Proficient with Terraform and IaC concepts
  • Hands-on Docker and CI/CD pipelines
  • Linux administration and troubleshooting
  • Python or Shell scripting for automation
  • Incident management, RCA and on-call experience
  • Solid networking fundamentals

Responsibilities

  • Maintain reliability, scalability, and observability of production systems
  • Manage AWS/GCP infrastructure and Kubernetes in production
  • Build and operate CI/CD pipelines and IaC tooling
  • Set up monitoring with Prometheus, Grafana, ELK/Loki and tracing
  • Handle on-call incidents, perform RCAs and post-incident reviews
  • Develop SLIs/SLOs, dashboards, and runbooks
  • Automate repetitive tasks using Python, Bash or Go
  • Plan capacity, autoscaling, performance improvements, and cost optimization
  • Troubleshoot Linux, networking, DNS, and load balancers
  • Drive preventive actions through automation and self-healing

Skills

SRE/DevOps
AWS
GCP
Kubernetes
Terraform
IaC
Docker
CI/CD
Linux
Python
Shell scripting
Networking basics
Incident management
On-call
OpenTelemetry

Tools

Kubernetes
Terraform
Docker
Prometheus
Grafana
ELK/Loki
OpenTelemetry

Job description

We are looking for an SRE to ensure the reliability, scalability, availability, and observability of high-traffic production systems. The role involves infrastructure automation, Kubernetes operations, monitoring, incident management, and continuous improvement of production reliability.

Key Responsibilities
  • Manage and troubleshoot AWS/GCP infrastructure and Kubernetes environments in production.
  • Work with Kubernetes, Redis, Kafka, Solr/Elasticsearch and related infrastructure components.
  • Build and maintain CI/CD pipelines, Terraform/Helm-based infrastructure, and automation.
  • Implement monitoring and alerting using Prometheus, Grafana, ELK/Loki and distributed tracing tools.
  • Participate in rotational on-call and shifts, handling P1/P2 incidents, troubleshooting, RCA, and post-incident reviews.
  • Develop SLIs/SLOs, alerts, dashboards, runbooks, and reliability improvements.
  • Automate repetitive operational tasks using Python, Shell/Bash, or Go to reduce manual toil.
  • Work on capacity planning, autoscaling, performance optimization, security patching, and cost optimization.
  • Troubleshoot Linux, networking, DNS, TCP/IP, load balancing, TLS/HTTPS, and application/infrastructure issues.
  • Drive preventive actions through RCA, automation, self-healing, and improved deployment/recovery processes.
Must-Have Skills
  • 3+ years of experience in SRE / DevOps / Infrastructure Engineering
  • Experience with high-traffic or large-scale production environments
  • Strong experience with AWS or GCP
  • Terraform and Infrastructure as Code (IaC)
  • Docker and CI/CD
  • Linux administration and troubleshooting
  • Python / Bash / Shell scripting
  • Production incident management, RCA and on-call experience
  • Good understanding of networking fundamentals
Good to Have
  • ELK/Loki
  • Redis, Kafka, Solr/Elasticsearch
  • SLI/SLO, SLA and error-budget concepts
  • Ansible
  • Distributed tracing / OpenTelemetry

Note: This is a rotational-shift/on-call role, so candidates should be comfortable supporting production systems across different shifts.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Snapmint • Gurugram District

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer (SRE) – DevOps Infrastructure
Site Reliability Engineer (SRE) – DevOps Infrastructure

PQAngels Technologies Pvt. Ltd. • Bengaluru

On-site
INR 1,200,000 - 2,100,000
Site Reliability Engineer
Site Reliability Engineer

Yantran • Chennai District

On-site
INR 900,000 - 1,400,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Smart Ims • Bengaluru

Hybrid
INR 1,200,000 - 2,000,000
SRE Engineer
SRE Engineer

ConsultBae India Private limited • India

On-site
INR 1,200,000 - 2,000,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

Skillventory • Kamrup Metropolitan

On-site
INR 3,500,000 - 7,000,000
Site Reliability Engineer (SRE) / DevOps Engineer
Site Reliability Engineer (SRE) / DevOps Engineer

New Era Technology • Gurugram District

On-site
INR 1,500,000 - 2,100,000
SRE Lead
SRE Lead

Acldigital • Ahmedabad District

On-site
INR 1,500,000 - 2,000,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Lonvec Technologies Private Limited • Hyderabad

On-site
INR 3,000,000 - 5,000,000