Senior Site Reliability Engineer

Carl Zeiss India (Bangalore) Pvt. Ltd.

Bengaluru

On-site

INR 2,500,000 - 4,500,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ZEISS India, headquartered in Bengaluru, is hiring a Site Reliability Engineer to bridge development, cloud platform engineering, and product owners across Digital Offerings.

The role focuses on defining service reliability (SLIs/SLOs), driving observability, building runbooks, and implementing automation and CI/CD pipelines. You will ensure availability, investigate incidents, and contribute to resilient, scalable cloud-based systems.

Qualifications

  • Define and measures the reliability of the service using SLI, SLOs and consider the risk minimization of service degradation.
  • Enable the development team to bring new software or new features to production quickly while ensuring IT operations performance and risk levels per SLAs.
  • Define and drive observability for self-developed software and managed cloud components by collecting observability data and setting up alerts.
  • Ensure availability and responsiveness of applications by documenting methods and tools.
  • Build Playbooks and Runbooks for troubleshooting to identify and investigate issues.
  • Set up automation and CI/CD pipelines for continuous deployment using best practices and review for improvements.
  • Monitor systems for performance and reliability, respond to incidents, and perform post-incident reviews for root cause analysis.

Responsibilities

  • Create and maintain documentation for systems, processes, recovery plans, and procedures.
  • Show expertise in system architecture, networking, and distributed systems.
  • Design and implement reliable, scalable, fault-tolerant systems.
  • Set up and manage monitoring, alerting, and logging for Kubernetes and related components (Prometheus, Grafana, OpenTelemetry).
  • Hands-on incident management including incident response, troubleshooting, and post-mortem analysis.
  • Use Terraform and other IaC tools for infrastructure automation and monitoring.
  • Stay updated with trends and technologies in site reliability engineering.
  • Collaborate with Software Developers, Product Owners, and Cloud Platform Engineers.

Skills

Observability
Incident management
CI/CD pipelines
Kubernetes
Prometheus
Grafana
Terraform
OpenTelemetry

Tools

Kubernetes
Prometheus
Grafana
OpenTelemetry
Terraform

Job description

ZEISS in India ZEISS in India is headquartered in Bengaluru and present in the fields of Industrial Quality Solutions, Research Microscopy Solutions, Medical Technology, Vision Care and Sports & Cine Optics. ZEISS India has 3 production facilities, R&D center, Global IT services and about 40 Sales & Service offices in almost all Tier I and Tier II cities in India. With 2200+ employees and continued investments over 25 years in India, ZEISS’ success story in India is continuing at a rapid pace. Further information at ZEISS India.

As a Site Reliability Engineer, you will bridge the gap between Development, Cloud Platform Engineering Team and Product Owners of different Digital Offerings.

What you will do:

Define and measures the reliability of the service using SLI, SLOs and consider the risk minimization of service degradation. Enable the development team to bring new software or new features to production as quickly as possible, while also ensuring an agreed-upon acceptable level of IT operations performance and error risk in line with the service level agreements (SLAs) agreed. Define and drive observability for self-developed software and the managed cloud components by collecting appropriate observability data for insights and alerting including setting up proper alerting for critical components. Ensure availability and responsiveness of application by setting up and maintaining the required documentation method and tools. Building Playbooks, Runbooks for troubleshooting techniques to effectively identify and investigate issues. Setup automation and CI/CD Pipelines for continuously deploying software applications using the best practices and review them for improvement. Monitor systems for performance and reliability, respond to incidents, and conduct post-incident reviews to identify root causes and improve system resilience. Create and maintain comprehensive documentation for systems, processes, disaster recovery plans, and procedures. In-depth knowledge of system architecture, networking, and distributed systems. Expertise in designing and implementing reliable, scalable, and fault-tolerant systems. Proficiency in setting up and managing monitoring, alerting, and logging systems for early detection and resolution of issues for container orchestrators like Kubernetes using Tools like Prometheus, Grafana, Open Telemetry Collector or similar tools. Hands-on experience in incident management, including incident response, troubleshooting, and post-mortem analysis. Proficiency in coding/scripting languages commonly used in infrastructure automation and monitoring (such as Terraform). Familiar with deployment process and strategies. Knowledge of best practices in disaster recovery planning and execution for cloud based Systems. Capability to advocate for SRE best practices and principles within the organization and drive cultural changes as needed. Willingness to stay updated with the latest trends, tools, and technologies in the field of site reliability engineering. Strong communication skills to effectively collaborate with cross-functional teams, including Software Developers, Product Owners, and Cloud Platform Engineers.

That’s just what our employees are doing every single day – in order to set the pace through our innovations and enable outstanding achievements. After all, behind every successful company are many great fascinating people. In a spacious modern setting full of opportunities for further development, ZEISS employees work in a place where expert knowledge and team spirit reign supreme. All of this is supported by a special ownership structure and the long‑term goal of the Carl Zeiss Foundation: to bring science and society into the future together.

Diversity is a part of ZEISS. We look forward to receiving your application regardless of gender, nationality, ethnic and social origin, religion, philosophy of life, disability, age, sexual orientation or identity.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Program Manager
Program Manager

Carl Zeiss India (Bangalore) Pvt. Ltd. • Bengaluru

On-site
INR 2,800,000 - 4,500,000
Site Reliability Engineer
Site Reliability Engineer

Foss United • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Customer Care Administrator
Customer Care Administrator

Carl Zeiss India (Bangalore) Pvt. Ltd. • New Delhi

On-site
INR 400,000 - 500,000
Customer Care Administrator
Customer Care Administrator

Carl Zeiss India (Bangalore) Pvt. Ltd. • Gurugram District

On-site
INR 300,000 - 600,000
Lead Developer- Salesforce
Lead Developer- Salesforce

Carl Zeiss India (Bangalore) Pvt. Ltd. • Bengaluru

On-site
INR 2,500,000 - 3,800,000
Marketing Specialist
Marketing Specialist

Carl Zeiss India (Bangalore) Pvt. Ltd. • Bengaluru

On-site
INR 1,200,000 - 1,600,000
Head of Customer Success, Last Mile Operations and Digitalization
Head of Customer Success, Last Mile Operations and Digitalization

Carl Zeiss India (Bangalore) Pvt. Ltd. • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Field Service Engineer
Field Service Engineer

ZEISS Group • Kolkata District

On-site
INR 600,000 - 1,000,000
Solution Architect - GenAI
Solution Architect - GenAI

Carl Zeiss India (Bangalore) Pvt. Ltd. • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Engineering Manager Linux and C++
Engineering Manager Linux and C++

ZEISS India • Bengaluru

On-site
INR 3,000,000 - 4,000,000