Site Reliability Engineer

Allegion India

Bengaluru

On-site

INR 1,800,000 - 3,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Allegion India is seeking a Site Reliability Engineer/DevOps to ensure reliability, scalability, and performance across a diverse product portfolio. You will design and maintain scalable infrastructure, collaborate with cross-functional teams, and implement automation to reduce manual work.

You will lead incident response, drive continuous improvement, and mentor junior engineers while participating in on-call rotations and DR planning.

Qualifications

  • Proven experience as a Site Reliability Engineer or similar role, focusing on highly available and scalable systems.
  • Cloud infrastructure on Azure with Docker and AKS; IaC using Terraform and ARM/Bicep.
  • Experience with AI/ML deployment and model serving endpoints on Azure ML or ML frameworks.
  • Advanced CI/CD pipelines with Azure DevOps or GitHub Actions.
  • Model observability and monitoring using Azure Monitor, Application Insights, Prometheus, Grafana, MLflow.
  • Security best practices on Azure, RBAC, cost optimization, AKS Spot, right-sizing.
  • Strong scripting in Python, Bash; solid networking knowledge.
  • Excellent communication, agile mindset, cross-functional collaboration.
  • Willingness to participate in on-call rotations and DR planning.

Responsibilities

  • Design, implement, and maintain highly available and scalable infrastructure systems to maximize uptime.
  • Collaborate with software teams to deploy applications using reliability and security best practices.
  • Develop automation tools to streamline operations and reduce manual work.
  • Monitor system performance, identify bottlenecks, and optimize for scalability.
  • Implement monitoring, alerting, and logging to detect issues before users are affected.
  • Lead incident response and root cause analysis for continuous improvement.
  • Define and enforce reliability standards and operational guidelines.
  • Participate in on-call rotations and disaster recovery planning.
  • Mentor junior engineers and promote learning and growth.
  • Run production environments with a holistic view of system health.

Skills

Cloud & Azure
AKS & Docker
Terraform
Bicep/ARM
ML Deployment
CI/CD
Azure DevOps
GitHub Actions
Model Monitoring
RBAC & Security
Scripting Python
Networking Basics
Agile
Communication

Education

BE/BTech/MTech/MCA/MSc in CS Engineering

Tools

Docker
Kubernetes
Prometheus
Grafana
MLflow
Seldon
KServe
Terraform
Ansible
CloudFormation
Azure DevOps
GitHub Actions

Job description

Allegion India is seeking a highly motivated Site Reliability Engineer or DevOps who will play a critical role in ensuring the reliability, scalability, and performance of our organization's systems and infrastructure, who will work with a team of cross-functional product development engineers to design, implement, and maintain highly available and resilient systems and whose expertise in automation, monitoring, and incident response will contribute to the overall stability and efficiency of our technology stack throughout the Allegion product portfolio.

What you’ll do:
  • Design, implement, and maintain highly available and scalable infrastructure systems, ensuring maximum uptime and performance.
  • Collaborate with software engineering teams to build and deploy applications using best practices in reliability, scalability, and security.
  • Develop and implement automation tools and frameworks to streamline operational processes, reduce manual intervention, and improve efficiency.
  • Monitor and analyse system performance, identifying bottlenecks, and implementing solutions to optimize performance and scalability.
  • Implement and maintain effective monitoring, alerting, and logging systems to proactively identify and resolve issues before they impact users.
  • Lead incident response and root cause analysis efforts, driving continuous improvement and preventing future incidents.
  • Collaborate with cross-functional teams to define and enforce best practices, standards, and guidelines for system reliability and performance.
  • Participate in on-call rotations and respond to incidents, ensuring timely resolution and minimal impact to users and thereby meeting SLAs.
  • Plan and devise Disaster Recovery (DR) strategies and implement DR Plans.
  • Mentor and provide guidance to junior team members, fostering a culture of learning and growth.
  • Run the production environment by monitoring availability and taking a holistic view of system health.
  • Build software and systems to manage platform infrastructure and applications.
  • Improve reliability, quality, and time-to-market of our suite of software solutions.
  • Measure and optimize system performance, with an eye toward pushing our capabilities forward, getting ahead of customer needs, and innovating for continual improvement.
  • Provide primary operational support and engineering for multiple large-scale distributed software applications.
What we are looking for:
Required Knowledge, Skills and Abilities:
  • Proven experience as a Site Reliability Engineer or similar role, with a focus on designing and maintaining highly available, reliable and scalable systems.
  • Cloud Infrastructure & Containerization: Design, provision, and manage scalable infrastructure on Microsoft Azure using containerization technologies (Docker) and orchestration platforms (Azure Kubernetes Service - AKS).
  • Infrastructure as Code (IaC): Automate the provisioning of secure and compliant environments (VNets, AKS clusters, Azure ML Workspaces, ACR) using Terraform and Bicep/ARM templates.
  • AI-Ops & Model Deployment: Partner with Data Scientists to productize machine learning models. Deploy scalable model serving endpoints (using Azure ML Managed Endpoints, KServe, or Seldon) with canary and blue-green rollout strategies.
  • Advanced CI/CD & Automation: Build robust CI/CD pipelines using Azure DevOps or GitHub Actions for both traditional application code and machine learning artifacts (automated model retraining and deployment).
  • Model Observability & Monitoring: Implement comprehensive monitoring for both system health and AI-specific metrics. Track model performance, data drift, and latency using Azure Monitor, Application Insights, Prometheus/Grafana, and MLflow. Elastic, Prometheus, Grafana, Loki, Sentry.io, CloudWatch, Wiz, etc.) for proactive system monitoring and troubleshooting.
  • Security & Cost Optimization: Implement Azure best practices for security (Managed Identities, RBAC) and optimize cloud costs by right-sizing compute, utilizing AKS Spot node pools, and managing batch inference schedules Proficient in configuration management tools like Ansible and infrastructure-as-code frameworks such as Terraform and CloudFormation.
  • Strong scripting skills (Python, Bash, etc.) to automate operational tasks and develop tooling.
  • Solid understanding of networking principles, protocols, and security best practices.
  • Strong problem-solving skills and the ability to work effectively in a fast‑paced, dynamic environment.
  • Excellent communication and collaboration skills, with the ability to work effectively with cross‑functional teams.
  • Proactive approach to identifying problems, performance bottlenecks, and areas for improvement.
  • Experience in Agile methodologies
  • Strong skills in software design, design patterns
  • Effective written, verbal and presentation skills with the ability to clearly articulate ideas and concepts.
  • Self‑directed and able to direct others.
Desired Skills & Abilities (Nice to have, but not required)
  • Experience with setting up performance/load test environments.
  • Familiarity with SOC2 audit processes

Experience : 3 to 6 Years of experience in DevOps/SRE/Software Application Development

Preferred Skills:
Mandatory Skills
  • Azure Cloud Platform: Deep, hands‑on experience managing infrastructure in Microsoft Azure, specifically with Azure Kubernetes Service (AKS), Azure Container Registry (ACR), and Azure Machine Learning (Azure ML).
  • Containerization & Orchestration: Production‑level expertise with Docker and Kubernetes. Ability to deploy, scale, and troubleshoot containerized applications and ML models on AKS.
  • YAML scripting
Desired Skills
  • Terraform
  • Python
  • Shell scripting
  • Wiz

Education : BE/BTech/M Tech/MCA/MSc in Computer Science Engineering

What we offer:

Allegion is a Great Place to Grow your Career if:

You are seeking a rewarding opportunity that allows you to truly help others. With thousands of employees and customers around the world, there’s plenty of room to make an impact. As our values state, “this is your business, run with it”.

  • You value personal well-being and balance because we do too!
  • You’re looking for a company that will invest in your professional development. As we grow, we want you to grow with us.
Work Culture:

Allegion is committed to building and maintaining a diverse and inclusive workplace. Together, we embrace all differences and similarities among colleagues, as well as the differences and similarities within the relationships that we foster with customers, suppliers, and the communities where we live and work.

Whatever your background, experience, religion, age, gender, gender identity, disability status, sexual orientation, or any other characteristic protected by law, we will make sure that you have every opportunity to impress us in your application and the opportunity to give your best at work, not because we’re required to, but because it’s the right thing to do.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead SRE
Lead SRE

United States Digital Space LLC • Karnataka

On-site
INR 900,000 - 1,400,000
Senior Devops Engineer
Senior Devops Engineer

Sitero • Bengaluru

Hybrid
INR 4,200,000 - 7,000,000
Health insurance
Hybrid work model
Learning budget
+1
Senior DevOps Engineer
Senior DevOps Engineer

Sitero LLC • Bengaluru

Hybrid
INR 3,000,000 - 5,000,000
Hybrid work model
Learning & development budget
Health insurance
+1
Lead SRE
Lead SRE

UST • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Staff Site Reliability Engineer - (Infra)
Staff Site Reliability Engineer - (Infra)

United States Digital Space LLC • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

United States Digital Space LLC • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000
Site Reliability Engineer II
Site Reliability Engineer II

United States Digital Space LLC • Bengaluru

Hybrid
INR 2,000,000 - 3,200,000
Staff Site Reliability Engineer - Ecosystem
Staff Site Reliability Engineer - Ecosystem

United States Digital Space LLC • Bengaluru

Hybrid
INR 4,000,000 - 6,000,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

BetterCloud • India

On-site
INR 1,500,000 - 2,500,000
Devops Architect
Devops Architect

ALSTOM Gruppe • Bengaluru

On-site
INR 1,400,000 - 1,800,000
Flexible and inclusive working environment
Investment in development and learning
Comprehensive social coverage (life, medical, pension)