Get more replies from employers
Send a job-specific resume in minutes.
Umanist Staffing LLC in Pune (Work From Office) is seeking an experienced Senior Site Reliability Engineer / DevOps Engineer to join our fast-paced team. The role requires deep expertise in Python/Bash, Azure cloud operations, OpenTelemetry, and Golden Signals with a strong SRE background.
You will own incident response, design scalable cloud infrastructure, and drive automation to reduce toil while maintaining high availability for critical production services.
Additional Important Note for Applicants
Important Note for Applicants
Kindly read the job description carefully before applying. Please apply only if your experience, technical skills, and notice period align with the mandatory requirements mentioned above. Profiles that do not meet the core criteria may face rejection during the screening process, which can lead to unnecessary time and effort from both sides. We appreciate your understanding and cooperation.
Location: Pune (Work From Office)
Experience: 10 Years
Shift Timing: 3:00 PM — 12:00 AM (Monday–Friday)
On-Call Requirement: 24/7 Production Support Rotation
Participate in 24/7 on-call rotation and production support.
Diagnose, troubleshoot, and resolve critical production incidents.
Lead Root Cause Analysis (RCA) and post-incident reviews.
Improve MTTR and overall operational efficiency.
Define and manage SLIs, SLOs, SLAs, and Error Budgets.
Drive reliability improvements, capacity planning, and disaster recovery readiness.
Reduce operational toil through automation and engineering solutions.
Design, implement, and manage cloud infrastructure on Microsoft Azure.
Manage Kubernetes clusters and containerized applications.
Implement Infrastructure as Code using Terraform.
Manage Helm deployments and Git-based CI/CD workflows.
Support highly available, scalable, and secure production environments.
Build and maintain monitoring and observability platforms.
Implement distributed tracing using OpenTelemetry.
Establish monitoring based on Golden Signals:
Design symptom-based alerting and proactive monitoring strategies.
Improve logging, tracing, metrics collection, and performance visibility.
Implement cloud security best practices.
Manage IAM, secrets management, and network security.
Support vulnerability remediation and compliance initiatives.
Must-Have SkillsExperience(Note: Candidates must have hands-on experience in Python/Bash, Azure Cloud Operations, OpenTelemetry, Golden Signals, and Site Reliability Engineering (SRE). Profiles lacking these mandatory skills should not be considered)
7 years in DevOps, Infrastructure Engineering Python or Bash programming/scripting, MS Azure Cloud Operations, Open Telemetery, Golden Signals and Site Reliability Engineering(mandate).
7 years of hands-on experience with DevOps tools and cloud-native infrastructure.
Experience supporting highly available production environments.
Python
Bash Scripting
Linux Administration
Networking Fundamentals (DNS, TCP/IP, Load Balancing, SSL/TLS)
Microsoft Azure (Mandatory)
Kubernetes
Terraform
Helm
GitHub / GitLab / Azure Repos
OpenTelemetry
Prometheus
Grafana
Datadog
Azure Monitor
Distributed Tracing
Metrics, Logs, and Observability Best Practices
Golden Signals Monitoring
Incident Response & Production Support
On-Call Operations
Root Cause Analysis (RCA)
SLI / SLO / SLA Management
Error Budgets
Capacity Planning
Reliability Engineering
Toil Reduction
AWS (EC2, S3, RDS, IAM, VPC, CloudWatch)
Google Cloud Platform (GCP)
Azure AI Services
AI Foundry
RAG (Retrieval-Augmented Generation) Infrastructure
AI/ML Production Workloads
OpenSearch
ELK Stack
Distributed Systems Architecture
Building observability frameworks using OpenTelemetry
Performance Engineering and System Optimization
Strong ownership mindset and accountability.
Excellent troubleshooting and debugging skills.
Experience handling critical production incidents calmly and effectively.
Deep understanding of SRE principles and operational excellence.
Strong collaboration and communication skills.
Passion for automation, scalability, and continuous improvement.