Site Reliability Engineer (SRE)

Zorba AI

Chennai District

On-site

INR 1,200,000 - 2,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zorba AI is seeking an experienced Site Reliability Engineer to join our engineering team. The role focuses on Azure/GCP, Kubernetes, IaC, CI/CD, monitoring, and production support in a microservices environment.

You will improve reliability, automate operations, and participate in incident management with on-call rotations. Strong troubleshooting and automation passion required.

Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
  • Strong experience with Azure and/or Google Cloud Platform (GCP).
  • Hands-on Kubernetes (AKS/GKE) administration.
  • Experience building CI/CD workflows with GitHub Actions.
  • IaC using Terraform and/or Ansible.
  • Experience with monitoring tools such as Splunk, Dynatrace, Grafana.
  • Familiarity with microservices architecture and reliability practices.
  • Knowledge of ITSM tools (e.g., ServiceNow) and ITIL processes.
  • Strong troubleshooting and production support capabilities.

Responsibilities

  • Implement SRE best practices to improve availability, scalability, and reliability.
  • Monitor production systems with proactive alerting and health checks.
  • Define and maintain SLI, SLO, and error budgets.
  • Participate in production incident management, RCA, and post-incident reviews.
  • Provide L2/L3 production support and on-call rotations.
  • Deploy and manage cloud infrastructure on Azure and/or GCP.
  • Automate provisioning and deployments using IaC and CI/CD pipelines.
  • Collaborate with development teams to improve release quality and efficiency.

Skills

Site Reliability Engineering
DevOps
Azure
GCP
Kubernetes
GitHub Actions
IaC
Terraform
Ansible
Monitoring tools
Splunk
Dynatrace
Grafana
Microservices
ITSM tools
ITIL
Troubleshooting
API-based architectures
Production support
On-call rotations

Tools

Terraform
Ansible
GitHub Actions
Splunk
Dynatrace
Grafana
AKS
GKE
ServiceNow
ITIL

Job description

Job Summary

We are seeking an experienced Site Reliability Engineer (SRE) to join our engineering team. The ideal candidate will have strong expertise in Azure/GCP cloud platforms, DevOps practices, Kubernetes, Infrastructure as Code, CI/CD, monitoring, and production support. The role focuses on improving application reliability, automating operations, optimizing system performance, and ensuring high availability for enterprise applications in a microservices environment.

The candidate should possess strong troubleshooting skills, experience with modern cloud-native architectures, and a passion for automation and continuous improvement.

Key Responsibilities Site Reliability Engineering
  • Implement Site Reliability Engineering (SRE) best practices to improve application availability, scalability, and reliability.
  • Monitor production systems and ensure service health through proactive monitoring and alerting.
  • Define and maintain SLI, SLO, and Error Budgets.
  • Participate in production incident management, root cause analysis (RCA), and post-incident reviews.
  • Provide L2/L3 production support and participate in on-call rotations.
Cloud Infrastructure
  • Deploy, manage, and maintain cloud infrastructure on Microsoft Azure and/or Google Cloud Platform (GCP).
  • Manage Kubernetes environments such as AKS and GKE.
  • Implement Infrastructure as Code (IaC) using Terraform and Ansible.
  • Automate cloud provisioning, deployments, and infrastructure management.
DevOps & CI/CD
  • Design and maintain CI/CD pipelines using GitHub Actions.
  • Implement automated build, testing, deployment, and release processes.
  • Improve deployment reliability through automation and DevOps best practices.
  • Collaborate with development teams to enhance release quality and deployment efficiency.
Monitoring & Reliability
  • Configure and maintain monitoring, logging, and alerting solutions.
  • Analyze application and infrastructure logs using Splunk, Dynatrace, Grafana, or similar tools.
  • Develop dashboards and monitoring metrics to ensure system reliability.
  • Improve monitoring through automation and preventive measures.
Automation
  • Develop automation scripts using Python or C#.
  • Automate operational tasks, deployments, housekeeping, and infrastructure management.
  • Improve operational efficiency by reducing manual interventions.
Production Support
  • Troubleshoot complex production issues across distributed systems.
  • Perform code analysis, log analysis, and performance tuning.
  • Collaborate with engineering teams to resolve critical production incidents.
  • Maintain production stability while minimizing downtime.
Collaboration
  • Work closely with cross-functional teams including Development, DevOps, QA, and Product teams.
  • Participate in Agile ceremonies and SRE governance activities.
  • Share operational best practices and contribute to continuous improvement initiatives.
Mandatory Skills
  • 5+ years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
  • Strong experience with Microsoft Azure and/or Google Cloud Platform (GCP).
  • Hands-on experience with Kubernetes (AKS/GKE).
  • Strong knowledge of DevOps practices and CI/CD pipelines.
  • Experience building CI/CD workflows using GitHub Actions.
  • Infrastructure as Code using Terraform and/or Ansible.
  • Experience with Splunk, Dynatrace, Grafana, or similar monitoring tools.
  • Strong experience with microservices architecture.
  • Experience with ServiceNow or other ITSM tools.
  • Knowledge of ITIL processes.
  • Strong troubleshooting and production support experience.
  • Experience working with SLI, SLO, Error Budget, and reliability engineering practices.
  • Strong understanding of API-based architectures.
  • Experience with desktop and mobile application support.

Skills: sre,devops,azure,reliability engineering

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE - AWS, GCP & Azure
SRE - AWS, GCP & Azure

PibyThree • Thane

On-site
INR 1,200,000 - 1,500,000
SRE - AWS, GCP & Azure
SRE - AWS, GCP & Azure

PibyThree • Navi Mumbai

On-site
INR 1,200,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

PwC India • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer (Azure) - S
Senior Site Reliability Engineer (Azure) - S

Tata Consultancy Services • Kolkata District, Chennai District, Bengaluru

On-site
INR 2,400,000 - 4,200,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Lonvec Technologies Private Limited • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Azure Site Reliability Engineer (SRE) - SaaS Operations
Azure Site Reliability Engineer (SRE) - SaaS Operations

Zensar • Pune District, Bengaluru

Hybrid
INR 1,800,000 - 2,400,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Augusta Infotech • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Tekskills • Hyderabad

Hybrid
INR 400,000 - 700,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Tekskills • Bulandshahr

Hybrid
INR 1,800,000 - 2,400,000