Site Reliability Manager

Staples India

Chennai District

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Staples India is seeking an experienced Manager of Information Technology to lead Site Reliability and Performance Engineering, supporting B2B and B2C sites with on-premises and public cloud. You will guide the Public Cloud Adoption initiative, mentor a cross-functional team, and drive MTTR/MTTD improvements with cutting-edge monitoring and testing tools.

The role requires 8+ years in SRE/performance roles, Azure expertise, and strong scripting and automation skills.

Qualifications

  • Bachelor’s degree in Computer Science or related field with 8+ years of experience.
  • Experience with Application Performance Management and monitoring tools.
  • Experience with on-premises and cloud deployments, especially Azure.
  • Strong scripting skills in Python or Azure CLI / PowerShell.
  • Hands-on with IaC and automation tools like Terraform, Puppet, Ansible, Jenkins.
  • Experience configuring and using monitoring and log analytics tools.
  • Familiar with load testing and performance tuning of web apps.
  • SRE/Cloud operations with 24x7 support expectations.

Responsibilities

  • Oversee day-to-day Site Reliability and Performance Engineering operations.
  • Set goals, supervise, and mentor the team.
  • Provide technical leadership and cross-functional collaboration.
  • Improve MTTR/MTTD with Product, Engineering, Security, Ops, Vendors.
  • Design, develop, and implement monitoring for availability and performance.
  • Plan and execute performance testing with JMeter, Locust, LoadRunner.
  • Automate performance testing in CI/CD pipelines.
  • Research and propose solutions to operational challenges.
  • Maintain knowledge documentation for the SRE team and partners.
  • Identify opportunities to innovate, automate, and reduce costs.

Skills

Python scripting
Azure CLI
PowerShell
Cloud technologies - Azure

Education

Bachelor’s degree in Computer Science or related field
Master’s degree in Computer Science or related field (preferred)

Tools

Terraform
Puppet
Ansible
Jenkins
New Relic
AppDynamics
SiteSpect
Datadog
Zabbix
Prometheus
JMeter
LoadRunner
Tomcat
Node.js
Spring Boot
Splunk
ELK/Elastic
Fullstory

Job description

Staples India Business Innovation Hub Private Limited | Permanent

As a Manager of Information Technology at Staples, you will collaborate with a business-critical team of engineers responsible for the B2B and B2C sites performance and availability of one of the top eCommerce companies in the United States. You will be a key contributor to the success of our Public Cloud Adoption initiative. This program will drive critical technology and tangible business value utilizing the latest cloud technologies. We are looking for a highly motivated and experienced Site Reliability and Performance Engineering leader who wants to grow their career and work with cutting-edge tools and technologies. The candidate must have a proven track record of supporting B2B, B2C sites and their integrations, both on-premises and in the public cloud, with demonstrated expertise in related technologies.

Duties & Responsibilities

Oversee the day-to-day operations of the Site Reliability and Performance Engineering team.

Set clear team goals, supervise, and manage the team.

Provide technical leadership and mentoring to team members.

Engage and collaborate with cross-functional Product, Engineering, Security, Operations, Infrastructure teams and Vendors to improve MTTD and MTTR

Design, develop, and implement infrastructure & application monitoring to ensure optimal platform availability and performance

Design and execute performance testing strategies including load, stress, and capacity planning using tools such as JMeter, Locust, and LoadRunner.

Automate performance testing within CI/CD pipelines to ensure continuous validation.

Research, analyze and recommend approaches for solving challenging operational issues

Develop and maintain robust knowledge documentation for the Site Reliability Engineering team and its partners

Proactively perform analysis and identify opportunities to innovate, automate, improve efficiency, and achieve cost savings

Foster innovation by encouraging new ideas and technologies within the team.

Ensure compliance with company standards and industry best practices.

Periodically review and assess the team's performance, providing feedback and facilitating professional growth.

Requirements
Basic Qualifications
  • Bachelor’s degree in Computer Science or related field with continuous and progressive experience
  • Minimum of 8 years of related experience working with these technologies:
  • Application Performance Management and Monitoring tools such as New Relic, AppDynamics, SiteSpect, and Datadog
  • Infrastructure monitoring tools like Zabbix, and Prometheus
  • Frameworks such as Dust/Angular, Nodejs, Springboot
  • Log Analytics tools like Splunk, and ELK/Elastic
  • Digital experience tools like Fullstory
  • Performance Testing tools such as JMeter, Loadrunner, etc.
  • Performance tuning experience with Tomcat, Node.js and Spring Boot.
  • Strong understanding of non-functional requirements, performance testing processes, and defect tracking.
  • 8+ years of experience with Cloud Technologies, at least half of which should be on the Microsoft Azure platform
  • Strong hands-on experience with infrastructure and services (systems, network, cloud technology, provisioning, storage, etc)
  • Must have strong experience with programming in one or more scripting languages (Python, Azure CLI, or Powershell)
  • Hands-on experience with tool sets related to automation, orchestration, and managing infrastructure (Terraform, Puppet, Ansible, or Jenkins)
  • Experience with configuring, deploying, and administering infrastructure and application monitoring tools that assist in troubleshooting performance and stability issues in a cloud environment.
  • As SRE and EIRE are global operational functions providing 24x7 support, weekend and public holiday coverage is an inherent expectation of these roles.
  • Eligible coverage will be offset through compensatory time off, aligned with company policy.
Preferred Qualifications

Master’s degree in Computer Science Software Engineering or a related field.

Certifications in project management or specific software development methodologies.

Experience in working with cross-functional teams and stakeholders at high organizational levels.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (EIRE)
Senior Site Reliability Engineer (EIRE)

Staples India • Chennai District

On-site
INR 1,200,000 - 2,200,000
Site Reliability Engineer
Site Reliability Engineer

Zorba AI • Hyderabad

On-site
INR 4,000,000 - 7,500,000
Site Reliability Engineer
Site Reliability Engineer

Zorba AI • Chennai District

On-site
INR 4,000,000 - 7,500,000
Tech & Digital-Site Reliability Engineer
Tech & Digital-Site Reliability Engineer

Hdfc Bank • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior Software Engineer - (Automation & Performance)
Senior Software Engineer - (Automation & Performance)

Staples India • Chennai District

On-site
INR 1,500,000 - 2,800,000
Software Engineer – Performance Engineering
Software Engineer – Performance Engineering

Staples India • Chennai District

On-site
INR 1,400,000 - 2,100,000
Resilience and Reliability Engineer
Resilience and Reliability Engineer

EY • Pune District, Gurugram District, Bengaluru

Hybrid
INR 1,800,000 - 2,800,000
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Maharashtra

On-site
INR 700,000 - 1,800,000
Lead SRE
Lead SRE

UST • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies • Bengaluru

On-site
INR 2,400,000 - 3,400,000