Site Reliability Engineer Lead

Good co India

United States

Remote

USD 120,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Good co India is seeking an experienced Site Reliability Engineer to lead and define SRE practices across AWS, Azure, GCP, and hybrid environments. The role focuses on incident management, performance optimization, and building scalable, fault-tolerant systems.

You will drive CI/CD, automation, and cost optimization while mentoring engineers and collaborating with cross-functional teams to improve reliability throughout the software lifecycle.

Qualifications

  • 5 to 10 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or related field.
  • Strong hands-on experience managing production systems and large-scale, distributed applications.
  • Strong proficiency in Linux and scripting/programming languages such as Python, Go, or Bash.
  • Extensive experience with AWS, Azure, GCP, or hybrid-cloud environments.
  • Strong experience with Kubernetes, Docker, Terraform, Ansible, and Infrastructure as Code.
  • Strong understanding of distributed systems, microservices, networking, databases, and cloud architecture.
  • Hands-on experience with observability platforms such as Prometheus, Grafana, ELK, Splunk, or OpenTelemetry.
  • Strong understanding of SLI, SLO, SLA, error budgets, availability, latency, reliability, and performance metrics.
  • Experience designing and operating highly available and fault-tolerant systems.
  • Strong experience in incident management, troubleshooting, root-cause analysis, and production support.
  • Experience building and managing CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar tools.
  • Knowledge of disaster recovery, business continuity, capacity planning, security, IAM, and cloud cost optimization.
  • Experience with automation, chaos engineering, performance engineering, and reducing operational toil is preferred.
  • Demonstrated ability to lead technical initiatives, mentor engineers, and work effectively across engineering teams.
  • Strong analytical, problem-solving, communication, and stakeholder-management skills.
  • Bachelors degree in Computer Science, Information Technology, Computer Engineering, or a related discipline; M.Tech/MS/MCA is an advantage.
  • Relevant certifications such as AWS/Azure/GCP, CKA/CKAD, Terraform, or ITIL are an advantage.

Responsibilities

  • Lead the Site Reliability Engineering function and ensure the availability, reliability, scalability, and performance of critical applications and infrastructure.
  • Define and implement SRE practices around SLIs, SLOs, SLAs, error budgets, and service reliability.
  • Design and operate highly available, fault-tolerant, and scalable cloud-native systems.
  • Lead incident response, troubleshooting, root-cause analysis, and post-incident reviews for critical production issues.
  • Build and improve monitoring, observability, logging, alerting, and performance-management capabilities.
  • Drive automation of infrastructure, deployment, operational, and repetitive engineering processes.
  • Manage and optimize cloud infrastructure across AWS, Azure, GCP, or hybrid environments.
  • Implement Infrastructure as Code using tools such as Terraform and Ansible.
  • Manage containerized and orchestration platforms including Docker and Kubernetes.
  • Establish and improve CI/CD pipelines, deployment strategies, release automation, and rollback mechanisms.
  • Lead capacity planning, performance optimization, disaster recovery, backup, and business-continuity initiatives.
  • Identify and reduce operational risks, technical debt, toil, and recurring production issues.
  • Implement reliability and resilience engineering practices, including failure testing and chaos engineering where appropriate.
  • Partner with development, infrastructure, security, and product teams to improve application reliability throughout the software lifecycle.
  • Establish SRE standards, operational runbooks, documentation, and engineering best practices.
  • Mentor SRE/DevOps engineers and provide technical leadership across reliability initiatives.
  • Track infrastructure and cloud costs and identify opportunities for performance and cost optimization.

Skills

SRE
DevOps
Cloud Engineering
Infrastructure Engineering
Distributed systems
Kubernetes
Docker
Terraform
Ansible
Python
Go
Bash
Linux
Prometheus
Grafana
ELK
OpenTelemetry
Jenkins
GitHub Actions
GitLab CI

Education

Bachelor's in Computer Science/IT/Computer Engineering
M.Tech/MS/MCA

Tools

Prometheus
Grafana
ELK
OpenTelemetry
GitHub Actions
Jenkins
GitLab CI
Terraform
Ansible
Docker
Kubernetes

Job description

Role & Responsibilities
  • Lead the Site Reliability Engineering function and ensure the availability, reliability, scalability, and performance of critical applications and infrastructure.
  • Define and implement SRE practices around SLIs, SLOs, SLAs, error budgets, and service reliability.
  • Design and operate highly available, fault-tolerant, and scalable cloud-native systems.
  • Lead incident response, troubleshooting, root-cause analysis, and post-incident reviews for critical production issues.
  • Build and improve monitoring, observability, logging, alerting, and performance-management capabilities.
  • Drive automation of infrastructure, deployment, operational, and repetitive engineering processes.
  • Manage and optimize cloud infrastructure across AWS, Azure, GCP, or hybrid environments.
  • Implement Infrastructure as Code using tools such as Terraform and Ansible.
  • Manage containerized and orchestration platforms including Docker and Kubernetes.
  • Establish and improve CI/CD pipelines, deployment strategies, release automation, and rollback mechanisms.
  • Lead capacity planning, performance optimization, disaster recovery, backup, and business-continuity initiatives.
  • Identify and reduce operational risks, technical debt, toil, and recurring production issues.
  • Implement reliability and resilience engineering practices, including failure testing and chaos engineering where appropriate.
  • Partner with development, infrastructure, security, and product teams to improve application reliability throughout the software lifecycle.
  • Establish SRE standards, operational runbooks, documentation, and engineering best practices.
  • Mentor SRE/DevOps engineers and provide technical leadership across reliability initiatives.
  • Track infrastructure and cloud costs and identify opportunities for performance and cost optimization.

Preferred Candidate Profile
  • 5 to 10 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or a related field.
  • Strong hands-on experience managing production systems and large-scale, distributed applications.
  • Strong proficiency in Linux and scripting/programming languages such as Python, Go, or Bash.
  • Extensive experience with AWS, Azure, GCP, or hybrid-cloud environments.
  • Strong experience with Kubernetes, Docker, Terraform, Ansible, and Infrastructure as Code.
  • Strong understanding of distributed systems, microservices, networking, databases, and cloud architecture.
  • Hands-on experience with observability platforms such as Prometheus, Grafana, ELK, Splunk, or OpenTelemetry.
  • Strong understanding of SLI, SLO, SLA, error budgets, availability, latency, reliability, and performance metrics.
  • Experience designing and operating highly available and fault-tolerant systems.
  • Strong experience in incident management, troubleshooting, root-cause analysis, and production support.
  • Experience building and managing CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar tools.
  • Knowledge of disaster recovery, business continuity, capacity planning, security, IAM, and cloud cost optimization.
  • Experience with automation, chaos engineering, performance engineering, and reducing operational toil is preferred.
  • Demonstrated ability to lead technical initiatives, mentor engineers, and work effectively across engineering teams.
  • Strong analytical, problem-solving, communication, and stakeholder-management skills.
  • Bachelors degree in Computer Science, Information Technology, Computer Engineering, or a related discipline; M.Tech/MS/MCA is an advantage.
  • Relevant certifications such as AWS/Azure/GCP, CKA/CKAD, Terraform, or ITIL are an advantage.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Gen Digital Inc. • United States

Remote
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Sr. Site Reliability Engineer(Local to Atlanta GA Only)
Sr. Site Reliability Engineer(Local to Atlanta GA Only)

Trigint Solutions LLC • Atlanta (GA)

Hybrid
USD 124,000 - 220,000
Site Reliability Engineer
Site Reliability Engineer

Moultrie • Birmingham (AL)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Site Reliability Engineering (SRE) Manager
Site Reliability Engineering (SRE) Manager

mtb • Buffalo (NY)

Hybrid
USD 150,000 - 230,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000