Senior Site Reliability Engineer

Jobgether SRL

India

Remote

INR 3,500,000 - 6,000,000

Full time

12 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Fully remote work from India
Remote-first collaboration across全球 팀
Technical mentorship and career growth
Opportunity to influence architecture

Job summary

Jobgether SRL in India is partnering with a client to hire a Senior Site Reliability Engineer who will design and operate highly available distributed systems at global scale. You will drive reliability across multi-region microservices on AWS and GCP, combining hands-on engineering with technical leadership.

You will define SLIs/SLOs, lead incident response, and push observability, automation, and FinOps improvements. This remote-first role offers mentorship and architectural influence.

Qualifications

  • 8+ years in SRE/DevOps for 24/7 distributed systems.
  • Bachelor’s degree or equivalent experience.
  • Advanced Python or Go, automation tooling, and API integrations.
  • Deep Kubernetes expertise in EKS/GKE, RBAC, networking, and GitOps.
  • Strong AWS/GCP knowledge and DR automation experience.

Responsibilities

  • Architect and scale reliability for multi-region microservices on AWS/GCP.
  • Lead disaster recovery, DR automation, and validation against RTO/RPO targets.
  • Define and enforce SLI/SLO and error-budget frameworks.
  • Own observability strategy with Datadog to reduce MTTR/MTTD.
  • Lead on-call escalation and post-incident reviews for lasting fixes.
  • Operate production Kubernetes environments and GitOps workflows (Argo CD).
  • Build Terraform IaC across multi-account, multi-region clouds.
  • Drive FinOps, cost optimization, and cloud spend visibility.
  • Mentor engineers and contribute to architectural standards.

Skills

Python/Go development
Kubernetes (EKS/GKE)
Terraform
AWS/GCP
SLI/SLO & observability
GitOps (Argo CD)
FinOps / cost optimization
Incident management / on-call
Chaos engineering

Education

Bachelor's degree in CS or equivalent

Tools

Datadog
Argo CD
Kargo
Istio/Linkerd
NGINX/HAProxy
Vault / AWS Secrets Manager

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in India.

This is a senior engineering role focused on building and operating highly available, resilient distributed systems at global scale.

You will architect reliability solutions across multi-region microservices, APIs, and authentication infrastructure running on AWS and GCP.

The role combines hands-on engineering with technical leadership across observability, disaster recovery, Kubernetes, and infrastructure automation.

You will help define reliability standards, including SLIs, SLOs, error budgets, and high-availability objectives.

You will also lead major incident response and turn production learnings into lasting systemic improvements.

Beyond reliability, you will drive cloud cost optimization, automation, and AI-assisted engineering practices.

This is an opportunity to mentor engineers, shape platform architecture, and raise the technical bar in a fast-moving, remote-first environment.

Accountabilities
  • Architect, scale, and continuously improve the reliability, availability, and performance of multi-region microservices, APIs, and authentication infrastructure across AWS and GCP.
  • Design and maintain disaster recovery and business continuity strategies, including multi-region failover automation, recovery dashboards, and validation processes aligned with defined RTO and RPO targets.
  • Establish and enforce SLI, SLO, and error-budget frameworks across engineering teams to improve reliability and accountability.
  • Lead the observability strategy using platforms such as Datadog, with actionable Golden Signals monitoring designed to reduce MTTD, MTTR, and alert fatigue.
  • Lead on-call escalation and major incident management, ensuring rapid response to production issues and strong adherence to availability objectives.
  • Facilitate blameless post-incident reviews and drive root-cause remediation to prevent recurring failures.
  • Architect and operate production Kubernetes environments, including EKS/GKE, networking, RBAC, ingress/egress, and GitOps workflows using tools such as Argo CD and Kargo.
  • Design and maintain modular, enterprise-grade Terraform infrastructure across complex multi-account and multi-region cloud environments.
  • Build FinOps dashboards and cost-optimization initiatives covering resource utilization, right-sizing, cost allocation, and multi-cloud spend visibility.
  • Develop production-grade Python or Go tooling, automation, integrations, and platform capabilities to eliminate operational toil.
  • Champion AI-assisted engineering tools to accelerate automation, runbook creation, incident triage, and engineering productivity.
  • Create operational runbooks and architecture documentation while mentoring junior and mid-level engineers and contributing to technical standards.
Requirements
  • 8+ years of professional experience in Site Reliability Engineering, DevOps, Platform Engineering, or software engineering supporting 24/7 mission-critical distributed systems.
  • Bachelor’s degree in Computer Science, Software Engineering, or a comparable technical discipline, or equivalent professional experience.
  • Advanced Python or Go development skills, with experience building internal SRE platforms, automation tools, and API integrations.
  • Deep hands-on Kubernetes expertise, including production EKS/GKE environments, cluster lifecycle management, networking, RBAC, ingress/egress, and GitOps tooling such as Argo CD.
  • Strong Terraform expertise, including module architecture, state management, refactoring, and infrastructure deployment across multi-account AWS environments.
  • Strong AWS and/or GCP expertise, including services and concepts such as IAM, VPC, Transit Gateway, ALB/NLB, Route53, and cloud networking.
  • Proven experience designing and testing multi-region disaster recovery architectures, automating failover, and monitoring recovery health.
  • Strong knowledge of SLI/SLO frameworks, production observability, PagerDuty, monitoring strategy, and reliability engineering practices.
  • Experience designing and operating enterprise service meshes such as Istio or Linkerd and production ingress/proxy technologies such as HAProxy or NGINX.
  • Demonstrated FinOps and cloud cost-optimization experience, including right-sizing, cost allocation, workload optimization, and financial visibility.
  • Experience with technical leadership, architectural discussions, RFCs/design documentation, and mentoring engineering peers.
  • Strong troubleshooting, communication, collaboration, and problem-solving skills, with the ability to work effectively on complex distributed systems.
  • Fluency in written and spoken English and the ability to collaborate with global teams.
  • Willingness and ability to participate in an on-call rotation and respond during assigned shifts.
  • Preferred experience includes chaos engineering, secrets management tools such as Vault or AWS Secrets Manager, DevSecOps practices, automated infrastructure vulnerability remediation, and identity or IAM-focused platforms.
Benefits
  • Fully remote work from India.
  • Remote-first working environment with collaboration across global teams.
  • Opportunity to work on highly available, mission-critical distributed systems at significant scale.
  • Exposure to modern cloud, Kubernetes, GitOps, Infrastructure-as-Code, observability, FinOps, and AI-assisted engineering technologies.
  • High level of technical ownership and autonomy in a fast-moving SaaS environment.
  • Opportunities to influence architecture, engineering practices, and reliability standards.
  • Technical mentorship and career development opportunities.
  • Collaborative culture that values connection, innovation, continuous improvement, and ambitious problem-solving.
  • Equal-opportunity workplace committed to an inclusive and diverse workforce.
  • Full-time employment with participation in an on-call rotation as part of the engineering role.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Lever, Inc. • India

Remote
INR 1,200,000 - 2,400,000
Fully remote in India
Global collaboration
AI-assisted tooling
+1
Site Reliability Engineer
Site Reliability Engineer

Jobgether SRL • India

Remote
INR 1,800,000 - 3,000,000
Fully remote in India
AWS/GCP exposure
Kubernetes & Terraform
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
AWS Delivery Manager
AWS Delivery Manager

EPAM Systems Inc • India

Hybrid
INR 6,000,000 - 9,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Staff S/W Engg , SRE & Platform Automation Engg
Staff S/W Engg , SRE & Platform Automation Engg

Aziro • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

Toyota Connected India • Chennai

On-site
INR 1,500,000 - 2,000,000
Yearly gym membership reimbursement
Free catered lunches
Flexible dress code
+1
Staff Software Engineer
Staff Software Engineer

Jobgether SRL • India

Remote
INR 1,300,000 - 2,100,000
Fully remote work from India
Ownership of modernization initiatives
Exposure to AI-assisted development
Site Reliability Engineer
Site Reliability Engineer

PwC Acceleration Center India • Bengaluru

On-site
INR 2,200,000 - 3,800,000
Staff Engineer (Core & MLOps)
Staff Engineer (Core & MLOps)

Lever, Inc. • India

Remote
INR 3,000,000 - 5,500,000
Fully remote
Remote-first culture
Flexible hours
+1