Site Reliability Engineer

Socure

Bengaluru

On-site

INR 1,500,000 - 2,100,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Socure is seeking Site Reliability Engineers to build, operate, and enhance production-grade systems in Bengaluru. You will work across AWS infrastructure, Kubernetes (EKS), automation, CI/CD, and observability with a strong emphasis on reliability and incident prevention.

You will own incident response, RCA, and long-term remediation, while improving monitoring, SLO adherence, and scalable CI/CD pipelines. Collaboration across engineering teams is essential for reliability improvements.

Qualifications

  • Hands-on experience with production systems.
  • Strong ownership and problem-solving capabilities.
  • Willingness to automate and simplify tasks.

Responsibilities

  • Operate and improve highly available AWS infrastructure.
  • Manage and troubleshoot Kubernetes (EKS) workloads and environments.
  • Improve production reliability through monitoring, alerting, automation, and SLOs.
  • Build and maintain reliable CI/CD pipelines.
  • Manage infrastructure using Terraform and GitOps.
  • Participate in on-call, incident response, RCA, and long-term remediation.
  • Automate repetitive operational tasks and reduce manual effort.
  • Partner with engineering teams to improve scalability, reliability, and operational readiness.

Skills

Troubleshooting
Production systems
Automation mindset
Incident response
Communication skills
Ownership mindset
Continuous learning

Tools

AWS
Kubernetes
Terraform
GitOps
CI/CD
ArgoCD
GitHub Actions
Jenkins
GitLab CI
Datadog
Prometheus
Grafana
CloudWatch

Job description

We are looking for Site Reliability Engineers who enjoy building, operating, and improving production-grade systems. You will work across AWS infrastructure, Kubernetes, automation, CI/CD, and observability, with a strong focus on improving reliability and preventing recurring production issues. This role is ideal for engineers who already have hands-on production experience and are looking to deepen their skills in cloud infrastructure, Kubernetes, automation, and reliability engineering.

Responsibilities:
  • Operate and improve highly available AWS infrastructure.
  • Manage and troubleshoot Kubernetes (EKS) workloads and environments.
  • Improve production reliability through monitoring, alerting, automation, and SLOs.
  • Build and maintain reliable CI/CD pipelines.
  • Manage infrastructure using Terraform and GitOps.
  • Participate in on-call, incident response, RCA, and long-term remediation.
  • Automate repetitive operational tasks and reduce manual effort.
  • Partner with engineering teams to improve scalability, reliability, and operational readiness.
Requirements:
  • Strong troubleshooting and problem-solving skills.
  • Experience working with real production systems.
  • Ownership mindset: willing to take a problem through to resolution.
  • Ability to learn from incidents and implement long-term fixes.
  • Strong bias toward automation and simplicity.
  • Curiosity about how systems work and why they fail.
  • Good communication and collaboration skills.
  • Willingness to continuously learn and improve.
Cloud and Infrastructure:
  • Good hands-on experience with AWS.
  • Understanding of networking, compute, IAM, scaling, and security fundamentals.
  • Experience managing infrastructure using Terraform.
Kubernetes and Platform Engineering:
  • Strong understanding of Kubernetes fundamentals.
  • Hands-on experience deploying and operating applications on Kubernetes.
  • Experience with Amazon EKS is preferred.
  • Ability to troubleshoot Kubernetes issues across pods, deployments, networking, resource utilization, and application health.
Coding and Automation:
  • Ability to write clean automation or scripts using Python or Go.
  • Strong automation mindset: look for opportunities to eliminate repetitive manual work.
CI/CD and GitOps:
  • Hands-on experience building or maintaining CI/CD pipelines.
  • Experience with GitHub Actions, Jenkins, GitLab CI, or similar tools.
  • Exposure to ArgoCD and GitOps-based deployment workflows.
Observability and Reliability:
  • Good understanding of metrics, logs, traces, monitoring, and alerting.
  • Experience with Datadog, Prometheus, Grafana, CloudWatch, or similar tools.
  • Understanding of application and infrastructure monitoring.
  • Familiarity with SLIs, SLOs, availability, MTTR, and incident management.
  • Ability to use monitoring and observability data to troubleshoot production issues.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Innodata Inc. • India

On-site
INR 2,400,000 - 4,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

UST • Pune District

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

DeepIQ • Hyderabad

On-site
INR 1,200,000 - 2,100,000
Site Reliability Engineer
Site Reliability Engineer

NOV • Ernakulam

On-site
INR 1,200,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

Smart Ims • Bengaluru

Hybrid
INR 1,200,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

Happiest Minds Technologies • Bengaluru

Hybrid
INR 2,400,000 - 3,600,000
Site Reliability Engineer
Site Reliability Engineer

Insight Global • Bengaluru

On-site
INR 2,500,000 - 5,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer Lead (Immediate Joiner)
Site Reliability Engineer Lead (Immediate Joiner)

HiLabs • Pune District

On-site
INR 2,500,000 - 4,200,000