Senior Site Reliability Engineer

Falabella India

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Falabella India is seeking an experienced Senior Site Reliability Engineer to join the Platform Engineering team in Bengaluru. You will design, build, and operate highly available cloud infrastructure, focusing on reliability, scalability, security, and cost efficiency.

You will own SLI/SLO definitions, drive automation, manage Kubernetes across cloud environments, and champion DevOps, GitOps, and DevSecOps practices to improve performance and developer productivity.

Qualifications

  • Bachelor's degree in computer science, engineering, or a related field.
  • 4+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering.
  • Strong experience managing production Kubernetes environments.
  • Hands-on expertise with public cloud platforms (AWS, GCP, or Azure).
  • Strong knowledge of Linux systems administration and networking.
  • Experience with Infrastructure as Code tools such as Terraform.
  • Expertise in observability and monitoring platforms.
  • Experience implementing CI/CD and GitOps practices.
  • Strong scripting and programming skills (Python, Go, Bash).
  • Experience with incident management and root cause analysis.

Responsibilities

  • Design, build, and operate highly available, scalable cloud infrastructure platforms.
  • Define SLIs, SLOs, and error budgets; drive reliability improvements.
  • Design, deploy, and manage large-scale Kubernetes environments across clouds.
  • Implement IaC with Terraform; manage Istio service mesh and cost optimisation.
  • Build monitoring, logging, and alerting platforms (Prometheus, Grafana, Datadog).
  • Lead incident response, on-call rotations, RCAs, and post-incident reviews.
  • Develop CI/CD and GitOps pipelines; enable self-service infra for engineers.
  • Collaborate with security to implement DevSecOps and compliance.

Job description

We are looking for an experienced Senior Site Reliability Engineer (SRE) to join our platform engineering team. The ideal candidate will be responsible for designing, building, and operating highly available, scalable, secure, and cost-efficient cloud infrastructure platforms. This role requires strong expertise in Kubernetes, cloud platforms, observability, automation, incident management, and reliability engineering practices. The Senior SRE will collaborate closely with software engineering, security, infrastructure, and operations teams to improve system reliability, performance scalability, and developer productivity.

Reliability Engineering
  • Design, implement, and maintain highly available and resilient production systems.
  • Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
  • Drive reliability improvements through automation and engineering best practices.
  • Perform capacity planning and performance optimisation for critical systems.
  • Conduct failure analysis and implement preventive measures.
Cloud Infrastructure And Kubernetes
  • Design, deploy, and manage large-scale Kubernetes environments.
  • Manage cloud infrastructure across AWS, GCP, or Azure environments.
  • Implement Infrastructure as Code (IaC) using Terraform.
  • Manage service mesh platforms such as Istio.
  • Optimise infrastructure utilisation and cloud costs.
Observability And Monitoring
  • Build and maintain monitoring, logging, and alerting platforms.
  • Implement observability solutions using tools such as Prometheus, Grafana, Datadog, VictoriaMetrics, and Loki.
  • Develop dashboards, alerts, and automated operational workflows.
  • Monitor system performance, availability, latency, and business metrics.
Incident Management And Operations
  • Lead production incident response and troubleshooting activities.
  • Participate in on-call rotations and major incident management.
  • Conduct root cause analysis (RCA) and post-incident reviews.
  • Drive continuous improvements to reduce operational toil.
  • Establish operational runbooks and automation frameworks.
Automation And DevOps
  • Develop automation scripts and tooling using Python, Go, or Bash.
  • Implement CI/CD pipelines and GitOps workflows.
  • Build self-service infrastructure capabilities for engineering teams.
  • Automate operational tasks and infrastructure provisioning.
Security And Compliance
  • Collaborate with security teams to implement DevSecOps practices.
  • Ensure infrastructure compliance with organisational standards.
  • Implement secure access controls, secrets management, and audit processes.
  • Support vulnerability management and remediation efforts.
FinOps And Optimisation
  • Monitor cloud spending and identify optimisation opportunities.
  • Implement rightsizing, reserved instance, and committed-use strategies.
  • Analyse infrastructure costs and recommend cost-saving initiatives.
  • Partner with engineering teams to improve resource efficiency.
Requirements
  • Bachelor's degree in computer science, engineering, or a related field.
  • 4+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering.
  • Strong experience managing production Kubernetes environments.
  • Hands-on expertise with public cloud platforms (AWS, GCP, or Azure).
  • Strong knowledge of Linux systems administration and networking.
  • Experience with Infrastructure as Code tools such as Terraform.
  • Expertise in observability and monitoring platforms.
  • Experience implementing CI/CD and GitOps practices.
  • Strong scripting and programming skills (Python, Go, Bash).
  • Experience with incident management and root cause analysis.
Preferred Qualifications
  • Experience operating large-scale distributed systems.
  • Experience with service mesh technologies (Istio, Linkerd).
  • Knowledge of database reliability engineering (PostgreSQL, MySQL, Redis).
  • Experience with security and compliance frameworks.
  • FinOps and cloud cost optimisation experience.
  • Experience managing multi-region and multi-cloud environments.
  • Kubernetes certifications (CKA, CKAD, CKS) are preferred.
Technical Skills
  • Cloud Platforms: AWS, Google Cloud Platform (GCP), Microsoft Azure.
  • Container and Platform: Kubernetes, Docker, Helm, Istio.
  • Infrastructure as Code: Terraform, Ansible.
  • Observability: Prometheus, Grafana, Datadog, Elasticsearch, VictoriaMetrics, Loki.
  • CI/CD and GitOps: GitHub Actions, GitLab CI, Jenkins, ArgoCD.
  • Databases: PostgreSQL, MySQL, Redis.
  • Programming: Python, Go, Bash.
Soft Skills
  • Strong troubleshooting and analytical skills.
  • Excellent communication and stakeholder management abilities.
  • Ability to lead complex technical initiatives.
  • Strong ownership mindset and operational excellence.
  • Ability to mentor engineers and drive engineering best practice.

This job was posted by Priyanka R N from Falabella.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

UST • Pune District

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

MishiPay • Bengaluru

On-site
INR 2,500,000 - 3,800,000
Senior DevOps / Site Reliability Engineer (SRE)
Senior DevOps / Site Reliability Engineer (SRE)

Aura Recruitment Solutions • Bengaluru

On-site
INR 450,000 - 750,000
SRE Engineer
SRE Engineer

Prodapt Solutions Private Limited • Chennai District

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

Epam Systems • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Socure • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

F-Prime Capital • Pune District

On-site
INR 1,500,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

NCR Voyix • Chennai District

On-site
INR 3,000,000 - 5,400,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

Innodata Inc. • India

On-site
INR 2,400,000 - 4,000,000