Site Reliability Engineer II

GreyOrange

Gurugram District

On-site

INR 3,000,000 - 5,500,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GreyOrange is seeking a Senior Site Reliability Engineer to strengthen our production systems, reduce outages, and improve platform reliability.

You will design scalable cloud infrastructure, automate deployments, and implement robust monitoring. You will work closely with software and ops teams to ensure service uptime and efficient incident response.

Qualifications

  • 5–8 years of experience in SRE or related roles.
  • Automation scripting in Python, Bash, or PowerShell for cloud environments.
  • Experience with observability tools: Grafana, Splunk, Dynatrace.
  • Automation/CM tools: Jenkins, GitLab, Ansible or Chef.
  • Containerization and orchestration with Docker and Kubernetes.
  • Understanding of SLI/SLO/SLA and error budgets.
  • On‑call support and incident management experience.
  • Troubleshooting production issues and bugs.
  • Strong Unix, networking, web tech, and database knowledge.
  • Cloud platform experience in AWS or GCP.

Responsibilities

  • Lead reliability projects and drive them to closure.
  • Ensure stability and high availability by monitoring performance.
  • Design and maintain scalable cloud-based infra and services.
  • Automate processes to improve observability and reduce toil.
  • Implement and manage observability for full monitoring and logging.
  • Own end-to-end availability of services and tools.
  • Support on-call incident response and postmortems.

Skills

Python
Bash
PowerShell
Observability
Incident management
Unix systems
Networking
AWS or GCP

Tools

Docker
Kubernetes
Jenkins
GitLab
Ansible
Chef
Grafana
Splunk
Dynatrace

Job description

We are seeking a talented and motivated Senior Site Reliability Engineer (SRE) to join our organization.

The SRE team at GreyOrange is responsible for monitoring the stability and availability of mission‑critical production systems, managing incidents for quicker resolution, and establishing BAU.

The team also manages and maintains internal tools/infra which is consumed by other development teams.

The experienced SRE will play a crucial role in ensuring the reliability, scalability, capacity planning, and performance of our infrastructure and applications.

The ideal candidate will have a strong background in software engineering, system administration, containerization, and cloud technologies.

REQUIREMENTS:
  • ? Should have 5 to 8 years of experience.
  • ? Well‑versed with scripting/programming languages (Python/Bash/PowerShell, etc.) to automate manual work, particularly within cloud environments.
  • ? Well‑versed with Observability tools (Grafana, Splunk, Dynatrace) for monitoring, alerting, and logging solutions to identify and assess potential issues, especially in cloud infrastructure.
  • ? Working experience with automation tools (Jenkins, GitLab, Ansible/Chef for configuration management) and processes to streamline deployment, monitoring, and management of systems and applications in the cloud.
  • ? Hands‑on experience with containerization and orchestration technologies such as Docker, Kubernetes, or similar, particularly in cloud‑native environments.
  • ? Well aware of SLI, SLO, SLA, and Error Budget concepts and their implementations.
  • ? Provide on‑call support and participate in incident management & response activities as needed.
  • ? Expert with troubleshooting production issues and bugs.
  • ? Good knowledge of Unix systems, networking, web technologies, and databases.
  • ? Incident Management experience coupled with effective communication skills for production workload.
  • ? Working knowledge in any one of the cloud platforms (AWS or GCP).
What you'll do:
  • ? Lead reliability engineering projects and drive them to closure.
  • ? Ensure system stability and high availability by proactively monitoring performance and troubleshooting issues.
  • ? Design, build and maintain efficient, reliable, and scalable cloud‑based infrastructure and services.
  • ? Automate processes and find opportunities to improve the observability and availability of the Platform to reduce toil.
  • ? Implement and manage observability tools for comprehensive monitoring, alerting, and logging.
  • ? Own end-to-end availability and performance of different services & tools.
  • ? Practice sustainable incident response and blameless postmortems.
  • ? Provide on‑call support for incident management and participate actively in response activities.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer 2
Site Reliability Engineer 2

GreyOrange • Gurugram District

Hybrid
INR 2,600,000 - 4,600,000
Site Reliability Engineer II
Site Reliability Engineer II

United States Digital Space LLC • Karnataka

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer
Site Reliability Engineer

Snapmint • Gurugram District

On-site
INR 800,000 - 1,200,000
Manager - SRE
Manager - SRE

GreyOrange • Gurugram District

On-site
INR 4,500,000 - 6,500,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Synechron • Bengaluru, Hyderabad

Hybrid
INR 4,200,000 - 6,300,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Keka Technologies Private Limited • Bengaluru

On-site
INR 4,000,000 - 7,000,000