Site Reliability Engineer

BuildxPartners

Sadar Bazar

On-site

INR 3,000,000 - 6,000,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Syfe is seeking an experienced Site Reliability Engineer/Platform Engineer to own the reliability of its production stack. You will manage a Kubernetes-native, multi-region platform across APAC regions, drive SLI/SLOs and incident response, and lead automation initiatives.

The role requires hands-on SRE experience with Kubernetes, cloud platforms (AWS), and observability tools, plus a strong ownership mindset and excellent communication during incidents.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 4–8 years of experience in SRE, Platform Engineering, or DevOps, with a strong senior IC track record of owning production systems.
  • Production-grade expertise with Kubernetes, containers, and cloud platforms (AWS preferred) in distributed-systems environments.
  • Hands-on experience defining and operating against SLIs, SLOs, and error budgets, with incident response and blameless postmortems.
  • Strong observability experience with metrics, logging, tracing, dashboards, and alerting (Datadog, Prometheus, Grafana).
  • Proficient in automation and infrastructure-as-code using Python, Go, Shell, and Terraform.
  • Comfortable with GitOps and CI/CD practices for reliable delivery.
  • Strong Linux/Unix fundamentals, networking, and cloud-native security knowledge.
  • Excellent communication skills for high-pressure incidents and written/oral updates.

Responsibilities

  • Own the end-to-end reliability of Syfe’s production platform, ensuring high availability and stability.
  • Operate a Kubernetes-native, multi-region platform across Singapore, Hong Kong, and Sydney.
  • Serve as a senior individual contributor, defining reliability standards and building systems and processes to meet them.
  • Define and drive SLIs, SLOs, and error budgets across critical services with product and engineering teams.
  • Own on-call, escalation, and incident-response programs including incident command and blameless postmortems.
  • Own the reliability of the AWS EKS-based deployment platform with GitOps (ArgoCD), Helm releases, and Terraform/OpenTofu; ensure safe, reversible deployments.
  • Build and continuously improve the observability stack (Datadog, Grafana, VictoriaMetrics, ClickHouse) and dashboards/alerts.
  • Lead capacity planning, scalability analysis, disaster recovery, and business continuity planning across regions.
  • Plan game days and chaos engineering to validate resilience and recovery processes.
  • Identify and reduce operational toil through automation and self-service tooling.

Skills

Kubernetes
AWS
GitOps
Terraform
Prometheus
Datadog
ArgoCD
CI/CD

Education

CS/Engineering degree

Tools

Terraform
OpenTofu
ClickHouse
VictoriaMetrics
Grafana
Datadog
Prometheus
ArgoCD

Job description

Sadar Bazar, India | Posted on 21/09/2026


BuildxPartners is a global talent solutions firm delivering end-to-end recruitment and workforce solutions across industries and geographies.


→ BuildxAlpha – Executive & Leadership Search Focused on C-suite, board, and global executive hiring.


→ BuildxSigma – Comprehensive Talent Across Levels Covering junior, mid-level, and senior professionals.


→ BuildxGCC – Global Capability Center Solutions Specializing in Build, Operate, Transfer (BOT) model for Global Capability Centers, GCC supports companies in setting up, scaling, and transferring GCCs.


Job Description

Qualifications



  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

  • 4–8 years of experience in SRE, Platform Engineering, or DevOps, with a strong senior IC track record of owning production systems.

  • Production-grade expertise with Kubernetes, containers, and cloud platforms (AWS preferred) in distributed-systems environments.

  • Hands‑on experience defining and operating against SLIs, SLOs, and error budgets, along with leading incident response and blameless postmortems.

  • Strong experience with observability tools covering metrics, logging, tracing, dashboards, and alerting, such as Datadog, Prometheus/Grafana, or equivalent platforms.

  • Proficient in automation and infrastructure‑as‑code, using technologies such as Python, Go, Shell, and Terraform.

  • Comfortable working with GitOps and CI/CD practices for reliable software and infrastructure delivery.

  • Strong understanding of Linux/Unix internals, networking, and cloud‑native security fundamentals.

  • Strong operational rigor, ownership mindset, and ability to communicate clearly in both written and verbal formats, including during high‑pressure incidents.


Responsibilities


  • Own the end-to-end reliability of Syfe’s production platform, ensuring high availability, performance, and operational stability.

  • Work with a Kubernetes-native, multi‑region platform across Singapore, Hong Kong, and Sydney, supporting a regulated digital wealth‑management product.

  • Operate as a senior individual contributor , defining measurable reliability standards and building systems, automation, and processes to maintain them.

  • Define and drive SLIs, SLOs, and error budgets across critical services, partnering with product and engineering teams to balance velocity and stability.

  • Own the on‑call, escalation, and incident‑response program , including incident command, blameless postmortems, RCA tracking, and reducing MTTD and MTTR.

  • Own the reliability of the AWS EKS‑based deployment platform , including GitOps with ArgoCD , Helm‑based release configuration, and Infrastructure as Code using Terraform/OpenTofu .

  • Ensure deployments are safe, progressive, and reversible , with strong rollout and rollback mechanisms.

  • Build and continuously improve the observability stack using tools such as Datadog, Grafana, VictoriaMetrics, and ClickHouse .

  • Develop meaningful dashboards, actionable alerts, and monitoring practices while reducing unnecessary alert noise.

  • Lead capacity planning, scalability analysis, failure‑mode analysis, disaster recovery, and business continuity planning across regions.

  • Plan and conduct game days and chaos engineering exercises to validate system resilience and recovery processes.

  • Identify operational toil and eliminate it through automation, self‑service tooling, and improved engineering practices .

  • Strengthen production safety , including deployment guardrails, rollback processes, and secrets management using HashiCorp Vault .

  • Partner with engineering teams to improve production readiness and service reliability before systems are deployed to production.

  • Drive reliability through hands‑on engineering, design reviews, production‑readiness reviews, runbooks, and technical documentation .

  • Mentor engineers and influence engineering teams to adopt SRE best practices and build a strong culture of production ownership.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - India
Senior Site Reliability Engineer - India

Syfe Pte. Ltd. • Gurgaon

On-site
INR 1,500,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BuildxPartners • Bengaluru Urban

Hybrid
INR 2,400,000 - 4,200,000
Lead Engineer - Reliability Engineering
Lead Engineer - Reliability Engineering

StoneX Group Inc. • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
SRE / Production Engineering
SRE / Production Engineering

Infosys • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Lead Site Reliability Engineer/ Expert
Lead Site Reliability Engineer/ Expert

SITA Group • Delhi

On-site
INR 1,200,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Seismic • Hyderabad

On-site
INR 4,000,000 - 7,500,000
Site Reliability Engineer
Site Reliability Engineer

Saika Technologies Inc. • Hyderabad, Bengaluru

Hybrid
INR 3,000,000 - 4,200,000
Delivery Lead-SRE
Delivery Lead-SRE

Acuity Analytics • Bengaluru

On-site
INR 1,400,000 - 2,400,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000