Senior Site Reliability Engineer (SRE) / DevOps Engineer

Umanist Staffing LLC

Maharashtra

On-site

INR 3,500,000 - 5,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Umanist Staffing LLC in Pune (Work From Office) is seeking an experienced Senior Site Reliability Engineer / DevOps Engineer to join our fast-paced team. The role requires deep expertise in Python/Bash, Azure cloud operations, OpenTelemetry, and Golden Signals with a strong SRE background.

You will own incident response, design scalable cloud infrastructure, and drive automation to reduce toil while maintaining high availability for critical production services.

Qualifications

  • Hands-on experience with Python or Bash in cloud operations.
  • Experience with Azure cloud operations and OpenTelemetry.
  • Strong background in SRE principles, incident management, and automation.

Responsibilities

  • Participate in 24/7 on-call rotation and production support.
  • Diagnose, troubleshoot, and resolve critical production incidents.
  • Lead Root Cause Analysis (RCA) and post-incident reviews.
  • Improve MTTR and overall operational efficiency.
  • Define and manage SLIs, SLOs, SLAs, and Error Budgets.
  • Drive reliability improvements, capacity planning, and disaster recovery readiness.
  • Reduce operational toil through automation and engineering solutions.

Skills

Python
Bash scripting
Azure Cloud Operations
OpenTelemetry
Golden Signals
Site Reliability Engineering
DevOps

Tools

Terraform
Kubernetes
Helm
GitHub
GitLab
Azure Repos
Prometheus
Grafana
Datadog
Azure Monitor
OpenTelemetry

Job description

Additional Important Note for Applicants

  • Currently, only immediate joiners (who have already completed their notice period) or candidates serving a notice period of up to 30 days will be considered for this opportunity.
  • Candidates with longer notice periods may not be considered at this stage due to urgent project requirements.

Important Note for Applicants

Kindly read the job description carefully before applying. Please apply only if your experience, technical skills, and notice period align with the mandatory requirements mentioned above. Profiles that do not meet the core criteria may face rejection during the screening process, which can lead to unnecessary time and effort from both sides. We appreciate your understanding and cooperation.

Senior Site Reliability Engineer (SRE) / DevOps Engineer

Location: Pune (Work From Office)

Experience: 10 Years

Shift Timing: 3:00 PM — 12:00 AM (Monday–Friday)

On-Call Requirement: 24/7 Production Support Rotation

Key ResponsibilitiesIncident Management & Reliability
  • Participate in 24/7 on-call rotation and production support.

  • Diagnose, troubleshoot, and resolve critical production incidents.

  • Lead Root Cause Analysis (RCA) and post-incident reviews.

  • Improve MTTR and overall operational efficiency.

  • Define and manage SLIs, SLOs, SLAs, and Error Budgets.

  • Drive reliability improvements, capacity planning, and disaster recovery readiness.

  • Reduce operational toil through automation and engineering solutions.

Cloud & Infrastructure
  • Design, implement, and manage cloud infrastructure on Microsoft Azure.

  • Manage Kubernetes clusters and containerized applications.

  • Implement Infrastructure as Code using Terraform.

  • Manage Helm deployments and Git-based CI/CD workflows.

  • Support highly available, scalable, and secure production environments.

Observability & Monitoring
  • Build and maintain monitoring and observability platforms.

  • Implement distributed tracing using OpenTelemetry.

  • Establish monitoring based on Golden Signals:

  • Latency
  • Traffic
  • Errors
  • Saturation
  • Design symptom-based alerting and proactive monitoring strategies.

  • Improve logging, tracing, metrics collection, and performance visibility.

Security & Compliance
  • Implement cloud security best practices.

  • Manage IAM, secrets management, and network security.

  • Support vulnerability remediation and compliance initiatives.

Must-Have SkillsExperience(Note: Candidates must have hands-on experience in Python/Bash, Azure Cloud Operations, OpenTelemetry, Golden Signals, and Site Reliability Engineering (SRE). Profiles lacking these mandatory skills should not be considered)

  • 7 years in DevOps, Infrastructure Engineering Python or Bash programming/scripting, MS Azure Cloud Operations, Open Telemetery, Golden Signals and Site Reliability Engineering(mandate).

  • 7 years of hands-on experience with DevOps tools and cloud-native infrastructure.

  • Experience supporting highly available production environments.

System & Programming Skills
  • Python

  • Bash Scripting

  • Linux Administration

  • Networking Fundamentals (DNS, TCP/IP, Load Balancing, SSL/TLS)

Cloud & Infrastructure
  • Microsoft Azure (Mandatory)

  • Kubernetes

  • Terraform

  • Helm

  • GitHub / GitLab / Azure Repos

Monitoring & Observability
  • OpenTelemetry

  • Prometheus

  • Grafana

  • Datadog

  • Azure Monitor

  • Distributed Tracing

  • Metrics, Logs, and Observability Best Practices

  • Golden Signals Monitoring

SRE Practices
  • Incident Response & Production Support

  • On-Call Operations

  • Root Cause Analysis (RCA)

  • SLI / SLO / SLA Management

  • Error Budgets

  • Capacity Planning

  • Reliability Engineering

  • Toil Reduction

Good-to-Have SkillsCloud Platforms
  • AWS (EC2, S3, RDS, IAM, VPC, CloudWatch)

  • Google Cloud Platform (GCP)

Programming
  • Go (Golang)
AI & Cloud-Native Workloads
  • Azure AI Services

  • AI Foundry

  • RAG (Retrieval-Augmented Generation) Infrastructure

  • AI/ML Production Workloads

Additional Technologies
  • OpenSearch

  • ELK Stack

  • Distributed Systems Architecture

Advanced Observability
  • Building observability frameworks using OpenTelemetry

  • Performance Engineering and System Optimization

Preferred Candidate Profile
  • Strong ownership mindset and accountability.

  • Excellent troubleshooting and debugging skills.

  • Experience handling critical production incidents calmly and effectively.

  • Deep understanding of SRE principles and operational excellence.

  • Strong collaboration and communication skills.

  • Passion for automation, scalability, and continuous improvement.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Maharashtra

On-site
INR 1,800,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

F-Prime Capital • Pune District

On-site
INR 1,500,000 - 2,000,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000
Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering
Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering

Deloitte & Touche GmbH Wirtschaftsprüfungsgesellschaft • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

New Era Technology • Gurugram District

On-site
INR 1,400,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
SRE - Site Reliability Engineering
SRE - Site Reliability Engineering

Build & Hire • Pune District

On-site
INR 1,500,000 - 2,300,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VMC Soft Technologies, Inc • Hyderabad

Hybrid
INR 1,500,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

ScaleneWorks People Solutions LLP • Pune District

On-site
INR 3,500,000 - 5,500,000