Site Reliability Engineer

Jobgether SRL

India

Remote

INR 1,800,000 - 3,000,000

Full time

11 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Fully remote in India
AWS/GCP exposure
Kubernetes & Terraform
GitOps with Argo CD
AI-assisted tooling
Disaster recovery focus

Job summary

Jobgether SRL is seeking a Site Reliability Engineer based in India for a fully remote, cloud-native role. You will design reliable production systems on AWS/GCP, build observability, and drive resilience through automation.

Key tasks include defining SLI/SLOs, incident management, and multi-cloud infra with Terraform. Strong English communication and collaboration with global teams are essential.

Qualifications

  • 5+ years in SRE/DevOps or platform engineering supporting 24/7 systems.
  • Hands-on Python or Go with automation experience.
  • Experience operating Kubernetes and GitOps with Argo CD.
  • Infra-as-code with Terraform in multi-cloud (AWS/GCP).

Responsibilities

  • Design, deploy, and maintain highly available production systems on AWS and GCP.
  • Define SLIs, SLOs, and error budgets with development teams.
  • Build end-to-end observability with Datadog and other tools.
  • Handle on-call rotations, incident response, and blameless post-incident reviews.
  • Manage EKS clusters, containers, and GitOps workflows (Argo CD).
  • Provision multi-cloud infra with Terraform; automate DR tests and runbooks.
  • Develop Python/Go automation to reduce toil and improve reliability.
  • Contribute to CI/CD, DevSecOps, and resilience testing initiatives.
  • Collaborate across distributed teams; communicate effectively in English.

Skills

Python
Go
Kubernetes
Argo CD
Terraform
AWS
GCP
Datadog
CI/CD
On-call
English
Python/Go automation
GitOps

Tools

Argo CD
Kubernetes
Terraform
GitHub Actions
GitLab Pipelines
Datadog
Istio
HAProxy
NGINX
HashiCorp Vault

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in India.

This is a fully remote engineering role focused on building resilient, highly available, and scalable production systems. You will operate at the intersection of software engineering, cloud infrastructure, observability, and platform reliability. The role involves designing automation, strengthening monitoring, reducing operational toil, and improving the reliability of cloud-native services. You will work extensively with AWS and GCP, Kubernetes, Terraform, and modern GitOps practices while supporting mission-critical systems. You will also contribute to disaster recovery, incident response, and reliability engineering standards across distributed services. This opportunity is well suited to an engineer who enjoys solving complex production problems through code and continuously improving operational maturity.

Accountabilities
  • Design, deploy, and maintain highly available, reliable, and performant production systems and APIs across AWS and GCP.
  • Define and operationalize SLIs, SLOs, and error budgets in partnership with application engineering teams.
  • Build and improve end-to-end observability across microservices and cloud infrastructure using Datadog or comparable monitoring platforms.
  • Implement actionable monitoring around the Golden Signals - latency, traffic, errors, and saturation to improve detection times and reduce unnecessary alerts.
  • Participate in on-call rotations, incident response, troubleshooting, and blameless post-incident reviews, using incidents to drive systemic improvements.
  • Manage production Kubernetes environments, including EKS clusters, container orchestration, and GitOps delivery workflows using tools such as Argo CD and Kargo.
  • Provision, manage, and secure multi-cloud infrastructure using modular Terraform and Infrastructure-as-Code practices.
  • Develop and maintain disaster recovery dashboards, runbooks, failover automation, and validation tests aligned with defined RTO and RPO targets.
  • Build production-grade Python or Go automation tools and scripts to eliminate repetitive operational work and reduce engineering toil.
  • Use AI-assisted development tools to accelerate scripting, runbook creation, troubleshooting, and incident triage.
  • Monitor cloud infrastructure costs, support resource optimization, and improve cost visibility through appropriate allocation and FinOps practices.
  • Configure and troubleshoot production service meshes and high-availability proxy solutions.
  • Contribute to CI/CD, DevSecOps, resilience testing, secrets management, and other platform engineering initiatives where required.
  • Proactively identify reliability, performance, security, and operational improvements across the production environment.
Requirements
  • 5+ years of professional software engineering experience in Site Reliability Engineering, DevOps, Platform Engineering, or a related discipline supporting 24/7 mission-critical systems.
  • Strong hands-on programming skills in Python or Go, with experience building SRE tools, automation, scripts, and cloud integrations.
  • Production experience operating Kubernetes clusters and containerized workloads, including experience with GitOps tools such as Argo CD.
  • Strong Infrastructure-as-Code experience using Terraform, including writing, maintaining, and modularizing configurations.
  • Direct experience operating cloud workloads on AWS or GCP, with knowledge of services such as EKS, IAM, VPC networking, Route 53, ALB/NLB, or comparable cloud infrastructure.
  • Practical FinOps experience, including cost-allocation tagging, resource right-sizing, cloud cost analysis, or development of cost-visibility dashboards.
  • Experience designing or operating disaster recovery solutions, conducting failover exercises, and monitoring recovery metrics.
  • Strong observability and incident-management experience using Datadog or similar tools, PagerDuty, alerting systems, and SLI/SLO frameworks.
  • Hands-on experience configuring and troubleshooting production service meshes such as Istio or equivalent technologies.
  • Experience managing high-availability proxy solutions such as HAProxy, NGINX, or comparable platforms.
  • Strong troubleshooting and problem-solving skills, with a demonstrated ability to improve operational efficiency through engineering and automation.
  • Experience with CI/CD platforms such as GitHub Actions or GitLab Pipelines is preferred.
  • Familiarity with chaos engineering or resilience testing in staging or production environments is a plus.
  • Knowledge of secrets-management solutions such as HashiCorp Vault, AWS Secrets Manager, or External Secrets Operator is beneficial.
  • Basic understanding of DevSecOps practices and Infrastructure-as-Code security scanning and remediation is preferred.
  • Strong collaboration and communication skills, with the ability to work effectively across distributed engineering teams.
  • Fluent written and spoken English, as the role involves working with global teams and conducting interviews and business communication primarily in English.
  • Willingness and ability to participate in an on-call rotation and respond during assigned shifts.
Benefits
  • Fully remote position for candidates based in India.
  • Opportunity to work on highly available, mission-critical cloud-native systems at scale.
  • Hands-on exposure to AWS, GCP, Kubernetes, Terraform, GitOps, observability, and modern reliability engineering practices.
  • Opportunity to use AI-assisted engineering tools to improve automation, incident response, and development productivity.
  • Significant scope to influence platform resilience, operational maturity, disaster recovery, and engineering efficiency.
  • Collaboration with experienced, globally distributed engineering teams.
  • Fast-paced SaaS environment offering challenging technical problems and opportunities for continuous learning.
  • Opportunity to contribute to reliability standards and engineering practices across distributed systems.
  • Supportive environment that values collaboration, innovation, continuous improvement, and technical ownership.

We appreciate your interest and wish you the best!

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Lever, Inc. • India

Remote
INR 1,200,000 - 2,400,000
Fully remote in India
Global collaboration
AI-assisted tooling
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Site Reliability Engineer 3 (Remote - India)
Site Reliability Engineer 3 (Remote - India)

Jobgether • India

On-site
INR 1,800,000 - 2,500,000
Competitive salary and bonuses
Flexible remote work
Paid time off and wellbeing days
+3
Site Reliability Engineer
Site Reliability Engineer

Acesoft Labs • Ahmedabad District

Hybrid
INR 400,000 - 700,000
Site Reliability Engineer
Site Reliability Engineer

Acesoft Labs • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Mumbai

On-site
INR 800,000 - 1,200,000
Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing • Pune District

On-site
INR 2,250,000 - 2,750,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Persistent Systems Limited • Pune District

On-site
INR 2,500,000 - 4,200,000
Competitive salary
Benefits package
Talent development
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hilabs • Bengaluru

On-site
INR 2,200,000 - 3,500,000