- Prometheus
- Linux & Networking
- Experience designing infrastructure for AI/ML platforms.
Location:
Viman Nagar, Pune – Work From Office
Experience Overall (must have)
8 Years
CTC
Up to ₹25 LPA
Notice Period
Immediate Joiners Only within 15d or (if serving max 30days)
Working Hours
3:00 PM – 12:00 AM, Monday to Friday
On-Call
24/7 Production Support – On-Call Rotation Required
Employment Type
Full-Time
About The Role
We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments. The role requires strong hands‑on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices. The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.
Must-Have Skills & Experience1. SRE & Production Operations
- Relevant 7+ years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering.
- Hands‑on experience with 24/7 production support and on‑call operations.
- Strong experience in incident management, troubleshooting, RCA, and post‑mortems.
- Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering.
- Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies.
- Ability to improve system availability, performance, scalability, and operational reliability.
- Cloud & Infrastructure
- Strong hands‑on experience with Microsoft Azure, AWS, and/or GCP.
- Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud‑native services.
- Hands‑on experience with at least one major cloud platform and good exposure to multi‑cloud environments.
- Experience with:
- Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS
- AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS
- GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring
- Kubernetes & Containerization
- Strong hands‑on experience with Kubernetes and containerized workloads.
- Experience with AKS / EKS / GKE or equivalent Kubernetes environments.
- Hands‑on experience with Helm deployments.
- Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.
- Infrastructure as Code & DevOps
- Hands‑on experience with Terraform / Infrastructure as Code (IaC).
- Experience with Git-based workflows using GitHub, GitLab, or Azure Repos.
- Strong DevOps automation and CI/CD understanding.
- Strong scripting skills in Python and or Bash.
- Monitoring & Observability
- Strong hands‑on experience with OpenTelemetry.
- Experience with monitoring and observability tools such as:
- Prometheus
- Grafana
- Datadog
- Azure Monitor
- AWS CloudWatch
- GCP Cloud Monitoring
- Strong understanding of metrics, logs, distributed tracing, and alerting.
- Experience implementing monitoring based on Golden Signals:
- Latency
- Traffic
- Errors
- Saturation
- Ability to develop symptom‑based, user‑impact‑focused alerting.
- Linux & Networking
- Strong knowledge of Linux system administration.
- Strong understanding of:
- DNS
- TCP/IP
- Load Balancing
- SSL/TLS
- Networking fundamentals
- Experience supporting highly available production environments.
- Incident & Reliability Engineering
- Ability to rapidly diagnose and resolve high‑severity production incidents.
- Experience driving MTTR reduction.
- Strong debugging and analytical problem‑solving skills.
- Ability to identify recurring issues and implement permanent corrective/preventive solutions.
Good-to-Have Skills
- Experience working across Azure + AWS + GCP in a multi‑cloud environment.
- Knowledge of Go (Golang).
- Experience with OpenSearch / ELK Stack.
- Experience supporting AI/ML workloads in production.
- Exposure to Azure AI Services and Azure AI Foundry.
- Experience supporting RAG (Retrieval-Augmented Generation) workloads.
- Experience designing infrastructure for AI/ML platforms.
- Experience building enterprise-wide OpenTelemetry observability frameworks.
- Strong understanding of distributed systems architecture.
- Exposure to advanced cloud‑native architectures and reliability patterns.
- Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.
Key ResponsibilitiesProduction & Incident Management
- Participate in the 24/7 on‑call rotation.
- Diagnose, mitigate, and resolve production incidents.
- Lead RCA and post‑incident reviews.
- Implement corrective and preventive actions.
- Continuously improve MTTR and production stability.
Reliability Engineering
- Define and improve SLIs, SLOs, SLAs, and Error Budgets.
- Identify and eliminate operational toil.
- Conduct reliability and capacity reviews.
- Improve redundancy, failover, disaster recovery, and system resilience.
Cloud & Infrastructure
- Manage and optimize cloud infrastructure across Azure, AWS, and or GCP.
- Manage Kubernetes clusters and containerized applications.
- Implement and maintain Infrastructure as Code using Terraform.
- Support CI/CD and Git-based development workflows.
Observability & Performance
- Build and improve monitoring, logging, metrics, and tracing.
- Implement OpenTelemetry and distributed tracing.
- Establish Golden Signals‑based monitoring and alerting.
- Identify and resolve infrastructure and application performance bottlenecks.
Security
- Implement cloud security best practices around IAM, network segmentation, and secrets management.
- Support vulnerability remediation and compliance initiatives.
- Collaborate with Development, Security, and Infrastructure teams.
Ideal Candidate
We Are Looking For Someone With
- Strong SRE mindset and production ownership.
- Excellent troubleshooting and incident‑management skills.
- Hands‑on expertise in Cloud + Kubernetes + Terraform + Observability.
- Strong understanding of OpenTelemetry and Golden Signals.
- Experience working in highly available, production‑critical environments.
- Ability to remain calm and make effective decisions during critical incidents.
- Strong communication and cross‑functional collaboration skills.
- Passion for automation, scalability, reliability, and continuous improvement.
Important Hiring Criteria
Must Be
- 7+ years relevant experience
- Immediate joiner
- Willing to work from office in Viman Nagar, Pune
- Comfortable with 3:00 PM – 12:00 AM shift
- Comfortable with 24/7 on-call rotation
- Strong hands‑on SRE/DevOps experience
- Strong Cloud + Kubernetes + Observability experience
- Strong production incident management experience
Good To Have
- Multi-cloud: Azure + AWS + GCP
- OpenTelemetry
- AI/ML or RAG production workloads
- Azure AI / AI Foundry
- Go
- OpenSearch / ELK
- Distributed systems
Skills: github,cloud infrastructure,linux system,gitlab,golden signals,ms azure,toil reduction,slo,aws,kubernetes & containerization,mttr,terraform,production engineering,iac,sre & production operations,error budgets,rca,troubleshooting,gcp,capacity planning,sla,24/7 production support,opentelemetry,devops,sli,python,golang,sre,incident management