Senior Associate Site Reliability Engineer

NTT DATA BUSINESS SOLUTIONS

Hyderabad

On-site

INR 1,400,000 - 2,000,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NTT DATA BUSINESS SOLUTIONS is seeking a Senior Associate Site Reliability Engineer to ensure reliability, availability, and performance of critical systems. You will monitor health, troubleshoot incidents, and collaborate with development and operations teams to improve resiliency and deployment processes.

You will work with automation and monitoring tools, participate in post-incident reviews, and help implement security and capacity planning practices while staying current with industry

Qualifications

  • Familiarity with cloud platforms (AWS/Azure/GCP) and Linux system administration.
  • Developing knowledge of programming or scripting languages and version control.
  • Understanding of incident management, monitoring tools and CM systems is beneficial.
  • Experience with performance monitoring and tuning tools (Prometheus, Grafana, New Relic).
  • Strong problem solving, communication and collaboration skills.

Responsibilities

  • Monitors system health, performance metrics, and alerts for incidents.
  • Diagnoses issues and restores services in a timely manner with teams.
  • Assists in deployment and release of software applications and infra changes.
  • Collaborates with development to minimize downtime during releases.
  • Automates routine tasks and improves operational efficiency.
  • Participates in capacity planning and scaling recommendations.
  • Documents incidents and post-incident reviews for preventive measures.
  • Works with security teams to implement best practices and compliance.

Skills

Cloud platforms
Scripting (Python/Bash)
Linux/Unix administration
Automation mindset
Collaboration & communication

Education

Bachelor's in CS/IT

Tools

Prometheus
Grafana
New Relic
Terraform
Jenkins
Git

Job description

Job Summary

The Senior Associate Site Reliability Engineer (SRE) is a developing subject matter expert responsible for playing a key role in ensuring the reliability, availability, and performance of company systems and infrastructure.

This role takes guidance from a supervisor and collaborates with cross functional teams to improve system resiliency and support the development and deployment of highly reliable software application.

The Senior Associate Site Reliability Engineer has opportunities to learn from experienced professionals, gain hands‑on experience, and grow their skills in ensuring the reliability and performance of critical systems and infrastructure.

Responsibilities
  • Monitors system health, performance metrics, and alerts to identify and respond to incidents promptly.
  • Works with Senior Site Reliability Engineers and teams to diagnose issues, troubleshoot problems, and restore services in a timely manner.
  • Assists in the deployment and release of software applications and infrastructure changes.
  • Collaborates with development teams to ensure smooth deployments, implement best practices, and minimize downtime during releases.
  • Collaborates with senior SREs and operations teams to automate routine tasks and improve operational efficiency.
  • Assists in capacity planning efforts, monitor resource utilization, and make recommendations for scaling infrastructure and services based on projected needs.
  • Collaborates with senior SREs to ensure adequate capacity to meet growing demands.
  • Documents incidents, their impact, and resolution procedures to maintain an incident knowledge base.
  • Participates in post-incident reviews, contributes to root cause analysis, and helps implement preventive measures to minimize future incidents.
  • Collaborates with security teams to implement security best practices and ensure compliance with industry standards and regulations.
  • Assists in monitoring and responding to security incidents, applying appropriate mitigation measures.
  • Works closely with development teams, operations teams, and other stakeholders to ensure smooth
  • Stays updated with the latest industry trends, emerging technologies, and best practices in Site Reliability Engineering.
  • Seeks opportunities to expand technical skills and knowledge through training, certifications, and self‑study.
  • Performs any other related task as required.
Requirements
  • Familiarity with infrastructure concepts, including cloud platforms (for example, AWS, Azure, Google Cloud), networking, and system administration.
  • Developing knowledge of programming or scripting languages (such as Python, Bash, or PowerShell) and version control systems (such as Git).
  • Relevant understanding of Linux/Unix systems and experience working with command‑line tools.
  • Strong problem‑solving and analytical skills, with attention to detail.
  • Excellent communication and collaboration skills, with the ability to work effectively in a team environment.
  • Passion for automation, reliability, and continuous improvement.
  • Familiarity with incident management processes, monitoring tools, and configuration management systems is beneficial.
  • Developing expertise in performance monitoring, optimization, and troubleshooting using tools such as Prometheus, Grafana, or New Relic.
  • Developing ability to optimize system performance, scalability, and reliability. experience with performance monitoring and tuning tools (for example, Prometheus, Grafana, or New Relic) to identify bottlenecks, analyze performance data, and implement optimization strategies.
  • Developing understanding of security principles, best practices, and compliance requirements.
Academic qualifications and certifications
  • Bachelor's degree or equivalent in Computer Science, Information Technology, or a related field.
  • Relevant certifications, such as AWS Certified DevOps Engineer - Professional, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator (CKA) preferred.
Required experience
  • Moderate level hands‑on experience in a Site Reliability Engineering role or related roles, including experience in designing and maintaining highly available and scalable systems.
  • Moderate level experience in incident response procedures and troubleshooting techniques to identify and resolve system issues.
  • Moderate level experience in automation principles and tools (for example, Terraform, Jenkins, Git).
  • Moderate level experience with scripting languages and version control systems, such as Git, demonstrates an ability to automate tasks and work collaboratively.

Workplace type: On-site

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Hyderabad

On-site
INR 1,800,000 - 3,200,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Augusta Infotech • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Lonvec Technologies Private Limited • Hyderabad

On-site
INR 3,000,000 - 5,000,000
SRE - AWS, GCP & Azure
SRE - AWS, GCP & Azure

PibyThree • Thane

On-site
INR 1,200,000 - 1,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Five9 • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

InOpTra Digital • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Resilience and Reliability Engineer
Resilience and Reliability Engineer

EY • Pune District, Gurugram District, Bengaluru

Hybrid
INR 1,800,000 - 2,800,000
SRE - AWS, GCP & Azure
SRE - AWS, GCP & Azure

PibyThree • Navi Mumbai

On-site
INR 1,200,000 - 1,800,000