Sadar Bazar, India | Posted on 21/09/2026
BuildxPartners is a global talent solutions firm delivering end-to-end recruitment and workforce solutions across industries and geographies.
→ BuildxAlpha – Executive & Leadership Search Focused on C-suite, board, and global executive hiring.
→ BuildxSigma – Comprehensive Talent Across Levels Covering junior, mid-level, and senior professionals.
→ BuildxGCC – Global Capability Center Solutions Specializing in Build, Operate, Transfer (BOT) model for Global Capability Centers, GCC supports companies in setting up, scaling, and transferring GCCs.
Job Description
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 4–8 years of experience in SRE, Platform Engineering, or DevOps, with a strong senior IC track record of owning production systems.
- Production-grade expertise with Kubernetes, containers, and cloud platforms (AWS preferred) in distributed-systems environments.
- Hands‑on experience defining and operating against SLIs, SLOs, and error budgets, along with leading incident response and blameless postmortems.
- Strong experience with observability tools covering metrics, logging, tracing, dashboards, and alerting, such as Datadog, Prometheus/Grafana, or equivalent platforms.
- Proficient in automation and infrastructure‑as‑code, using technologies such as Python, Go, Shell, and Terraform.
- Comfortable working with GitOps and CI/CD practices for reliable software and infrastructure delivery.
- Strong understanding of Linux/Unix internals, networking, and cloud‑native security fundamentals.
- Strong operational rigor, ownership mindset, and ability to communicate clearly in both written and verbal formats, including during high‑pressure incidents.
Responsibilities
- Own the end-to-end reliability of Syfe’s production platform, ensuring high availability, performance, and operational stability.
- Work with a Kubernetes-native, multi‑region platform across Singapore, Hong Kong, and Sydney, supporting a regulated digital wealth‑management product.
- Operate as a senior individual contributor , defining measurable reliability standards and building systems, automation, and processes to maintain them.
- Define and drive SLIs, SLOs, and error budgets across critical services, partnering with product and engineering teams to balance velocity and stability.
- Own the on‑call, escalation, and incident‑response program , including incident command, blameless postmortems, RCA tracking, and reducing MTTD and MTTR.
- Own the reliability of the AWS EKS‑based deployment platform , including GitOps with ArgoCD , Helm‑based release configuration, and Infrastructure as Code using Terraform/OpenTofu .
- Ensure deployments are safe, progressive, and reversible , with strong rollout and rollback mechanisms.
- Build and continuously improve the observability stack using tools such as Datadog, Grafana, VictoriaMetrics, and ClickHouse .
- Develop meaningful dashboards, actionable alerts, and monitoring practices while reducing unnecessary alert noise.
- Lead capacity planning, scalability analysis, failure‑mode analysis, disaster recovery, and business continuity planning across regions.
- Plan and conduct game days and chaos engineering exercises to validate system resilience and recovery processes.
- Identify operational toil and eliminate it through automation, self‑service tooling, and improved engineering practices .
- Strengthen production safety , including deployment guardrails, rollback processes, and secrets management using HashiCorp Vault .
- Partner with engineering teams to improve production readiness and service reliability before systems are deployed to production.
- Drive reliability through hands‑on engineering, design reviews, production‑readiness reviews, runbooks, and technical documentation .
- Mentor engineers and influence engineering teams to adopt SRE best practices and build a strong culture of production ownership.