Site Reliability Engineer

Lever, Inc.

India

Remote

INR 1,200,000 - 2,400,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Fully remote in India
Global collaboration
AI-assisted tooling
SaaS environment growth

Job summary

Lever, Inc. is seeking a Site Reliability Engineer based in India for a fully remote role focused on building resilient, scalable production systems across AWS and GCP.

You will design automation, observability, and disaster recovery, while reducing operational toil and enhancing platform reliability. You will work with Kubernetes, Terraform, and GitOps (Argo CD) and participate in on-call rotations, incident response, and cost-visibility initiatives.

Qualifications

  • 5+ years of professional software engineering in SRE/DevOps/Platform Engineering supporting 24/7 systems.
  • Strong Python or Go programming with automation experience.
  • Experience operating Kubernetes clusters and GitOps tools (Argo CD).
  • Terraform IaC experience, modular configurations for multi-cloud.
  • Hands-on AWS or GCP cloud workloads (EKS, IAM, VPC, Route 53, ALB/NLB).
  • FinOps exposure with cost-visibility dashboards and tagging.

Responsibilities

  • Design, deploy, and maintain highly available production systems across AWS and GCP.
  • Define SLIs/SLOs and error budgets with engineering teams.
  • Build end-to-end observability with Datadog and alerting systems.
  • Manage on-call rotation, incident response, and blameless post-incident reviews.
  • Operate Kubernetes environments and GitOps workflows via Argo CD and related tools.
  • Develop automation in Python/Go to reduce toil and improve efficiency.
  • Contribute to CI/CD, DevSecOps, and resilience testing initiatives.

Skills

Python/Go programming
Cloud infrastructure
Observability
SRE tooling
CI/CD automation
FinOps
Disaster recovery planning
Incident management
On-call readiness
Team collaboration
English communication

Tools

Kubernetes
Terraform
Argo CD
GitHub Actions
GitLab Pipelines
Istio
NGINX
HAProxy
Datadog
PagerDuty
AWS
GCP
HashiCorp Vault
AWS Secrets Manager

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in India.

This is a fully remote engineering role focused on building resilient, highly available, and scalable production systems. You will operate at the intersection of software engineering, cloud infrastructure, observability, and platform reliability. The role involves designing automation, strengthening monitoring, reducing operational toil, and improving the reliability of cloud-native services. You will work extensively with AWS and GCP, Kubernetes, Terraform, and modern GitOps practices while supporting mission-critical systems. You will also contribute to disaster recovery, incident response, and reliability engineering standards across distributed services. This opportunity is well suited to an engineer who enjoys solving complex production problems through code and continuously improving operational maturity.

Accountabilities
  • Design, deploy, and maintain highly available, reliable, and performant production systems and APIs across AWS and GCP.

  • Define and operationalize SLIs, SLOs, and error budgets in partnership with application engineering teams.

  • Build and improve end-to-end observability across microservices and cloud infrastructure using Datadog or comparable monitoring platforms.

  • Implement actionable monitoring around the Golden Signals-latency, traffic, errors, and saturation to improve detection times and reduce unnecessary alerts.

  • Participate in on-call rotations, incident response, troubleshooting, and blameless post-incident reviews, using incidents to drive systemic improvements.

  • Manage production Kubernetes environments, including EKS clusters, container orchestration, and GitOps delivery workflows using tools such as Argo CD and Kargo.

  • Provision, manage, and secure multi-cloud infrastructure using modular Terraform and Infrastructure-as-Code practices.

  • Develop and maintain disaster recovery dashboards, runbooks, failover automation, and validation tests aligned with defined RTO and RPO targets.

  • Build production-grade Python or Go automation tools and scripts to eliminate repetitive operational work and reduce engineering toil.

  • Use AI-assisted development tools to accelerate scripting, runbook creation, troubleshooting, and incident triage.

  • Monitor cloud infrastructure costs, support resource optimization, and improve cost visibility through appropriate allocation and FinOps practices.

  • Configure and troubleshoot production service meshes and high-availability proxy solutions.

  • Contribute to CI/CD, DevSecOps, resilience testing, secrets management, and other platform engineering initiatives where required.

  • Proactively identify reliability, performance, security, and operational improvements across the production environment.

Requirements
  • 5+ years of professional software engineering experience in Site Reliability Engineering, DevOps, Platform Engineering, or a related discipline supporting 24/7 mission-critical systems.

  • Strong hands-on programming skills in Python or Go, with experience building SRE tools, automation, scripts, and cloud integrations.

  • Production experience operating Kubernetes clusters and containerized workloads, including experience with GitOps tools such as Argo CD.

  • Strong Infrastructure-as-Code experience using Terraform, including writing, maintaining, and modularizing configurations.

  • Direct experience operating cloud workloads on AWS or GCP, with knowledge of services such as EKS, IAM, VPC networking, Route 53, ALB/NLB, or comparable cloud infrastructure.

  • Practical FinOps experience, including cost-allocation tagging, resource right-sizing, cloud cost analysis, or development of cost-visibility dashboards.

  • Experience designing or operating disaster recovery solutions, conducting failover exercises, and monitoring recovery metrics.

  • Strong observability and incident-management experience using Datadog or similar tools, PagerDuty, alerting systems, and SLI/SLO frameworks.

  • Hands-on experience configuring and troubleshooting production service meshes such as Istio or equivalent technologies.

  • Experience managing high-availability proxy solutions such as HAProxy, NGINX, or comparable platforms.

  • Strong troubleshooting and problem-solving skills, with a demonstrated ability to improve operational efficiency through engineering and automation.

  • Experience with CI/CD platforms such as GitHub Actions or GitLab Pipelines is preferred.

  • Familiarity with chaos engineering or resilience testing in staging or production environments is a plus.

  • Knowledge of secrets-management solutions such as HashiCorp Vault, AWS Secrets Manager, or External Secrets Operator is beneficial.

  • Basic understanding of DevSecOps practices and Infrastructure-as-Code security scanning and remediation is preferred.

  • Strong collaboration and communication skills, with the ability to work effectively across distributed engineering teams.

  • Fluent written and spoken English, as the role involves working with global teams and conducting interviews and business communication primarily in English.

  • Willingness and ability to participate in an on-call rotation and respond during assigned shifts.

Benefits
  • Fully remote position for candidates based in India.
  • Opportunity to work on highly available, mission-critical cloud-native systems at scale.
  • Hands-on exposure to AWS, GCP, Kubernetes, Terraform, GitOps, observability, and modern reliability engineering practices.
  • Opportunity to use AI-assisted engineering tools to improve automation, incident response, and development productivity.
  • Significant scope to influence platform resilience, operational maturity, disaster recovery, and engineering efficiency.
  • Collaboration with experienced, globally distributed engineering teams.
  • Fast-paced SaaS environment offering challenging technical problems and opportunities for continuous learning.
  • Opportunity to contribute to reliability standards and engineering practices across distributed systems.
  • Supportive environment that values collaboration, innovation, continuous improvement, and technical ownership.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing • Pune District

On-site
INR 2,250,000 - 2,750,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer 3 (Remote - India)
Site Reliability Engineer 3 (Remote - India)

Jobgether • India

On-site
INR 1,800,000 - 2,500,000
Competitive salary and bonuses
Flexible remote work
Paid time off and wellbeing days
+3
Site Reliability Engineer
Site Reliability Engineer

Acesoft Labs • Ahmedabad District

Hybrid
INR 400,000 - 700,000
Site Reliability Engineer
Site Reliability Engineer

Acesoft Labs • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SourcingXPress • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Persistent Systems Limited • Pune District

On-site
INR 2,500,000 - 4,200,000
Competitive salary
Benefits package
Talent development
+4
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000