Site Reliability Engineer

NTT DATA BUSINESS SOLUTIONS

Hyderabad

On-site

INR 1,800,000 - 3,200,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NTT DATA is seeking a seasoned Site Reliability Engineer (SRE) in Hyderabad to ensure reliability, availability, and performance of critical systems. You will partner with development and operations teams to enhance resiliency, automate processes, and implement scalable solutions across on‑prem and cloud environments.

You will design and maintain automation tools, implement infrastructure-as-code, and lead incident response efforts.

Qualifications

  • Seasoned Linux/Unix administration and network knowledge.
  • Experience with cloud platforms (AWS/Azure/GCP) and services.
  • Proficient in at least one programming language (Python/Go/Java/Ruby).
  • Strong monitoring, performance tuning, and incident management experience.
  • Knowledge of IaC (Terraform/CloudFormation) and containerization (Docker/Kubernetes).
  • Ability to design robust automation, CI/CD pipelines, and deployment strategies.
  • Understanding of security best practices and compliance requirements.

Responsibilities

  • Monitors system health, performance metrics, and alerts; promptly responds to incidents.
  • Implements incident response processes to minimize downtime and boost availability.
  • Designs and maintains automation tools, scripts, and deployment processes.
  • Applies infrastructure-as-code to ensure consistency and repeatability.
  • Optimizes resources for performance and scalability across environments.
  • Uses monitoring tools to identify bottlenecks and optimize systems.
  • Works with teams to forecast capacity needs and plan growth.
  • Leads incident response, coordinating across teams and driving resolution.
  • Performs post-incident analysis to identify root causes and preventive measures.
  • Drives automation, self-healing, monitoring, and deployment framework initiatives.
  • Aims to improve operational efficiency and consistency across environments.

Skills

Linux/Unix
Python
Go
Java
Ruby
Bash
Monitoring & alerting
Cloud platforms
Incident management
CI/CD
Incident response

Education

Bachelor's degree in Computer Science / IT
AWS DevOps Engineer – Professional
Google Cloud Professional DevOps Engineer
Certified Kubernetes Administrator (CKA)

Tools

Terraform
CloudFormation
Docker
Kubernetes
Jenkins
GitLab CI/CD
CircleCI
Prometheus
Grafana
New Relic

Job description

Job Summary

Make an impact with NTT DATA Join a company that is pushing the boundaries of what is possible. We are renowned for our technical excellence and leading innovations, and for making a difference to our clients and society. Our workplace embraces diversity and inclusion its a place where you can grow, belong and thrive.

Your day at NTT DATA

The Site Reliability Engineer (SRE) is a seasoned subject matter expert, responsible for ensuring the reliability, availability, and performance of company systems and infrastructure. The Site Reliability Engineer (SRE) works closely with development teams, operations teams, and other stakeholders to enhance system resiliency, automate processes, and improve overall system reliability.

Responsibilities
  • Monitors system health, performance metrics, and alerts to identify and respond to incidents promptly and diagnoses issues, troubleshoots problems, and restores services in a timely manner.
  • Implements incident response processes to minimize downtime and improve system availability.
  • Designs, develops, and maintains automation tools, scripts, and processes to streamline system management tasks, deployments, and configuration changes.
  • Implements infrastructure-as-code principles to ensure consistency and repeatability.
  • Optimizes system resources, configurations, and processes to enhance performance, scalability, and efficiency.
  • Uses monitoring tools and performance testing to identify bottlenecks and implement optimizations.
  • Collaborates with teams to forecast system resource needs, plans for capacity growth, and ensures adequate scalability.
  • Leads incident response efforts, coordinates with cross-functional teams, and drives the resolution of system issues.
  • Performs thorough post-incident analysis to identify root causes and implements preventive measures to minimize future incidents.
  • Identifies opportunities for automation and drives the implementation of self-healing, monitoring, and deployment of automation tools and frameworks.
  • Continuously improves operational efficiency, system reliability, and availability through process enhancements and automation.
  • Ensures consistency across environments, tracks changes, and enforces configuration standards.
  • Works closely with development teams, operations teams, and other stakeholders to ensure effective collaboration, knowledge sharing, and alignment on reliability goals.
  • Implements security best practices, works with security teams to assess and address vulnerabilities, and ensures compliance with security standards and regulations.
  • Performs any other related task as required.
Requirements
  • Seasoned technical expertise in Linux/Unix systems, networking, and system administration.
  • Seasoned proficiency in scripting or programming languages, such as Python, Go, Java, or Ruby.
  • Seasoned knowledge of cloud platforms (such as AWS, Azure, or Google Cloud) and associated services.
  • Seasoned proven expertise in performance monitoring, optimization, and troubleshooting using tools such as Prometheus, Grafana, or New Relic.
  • Seasoned expertise in incident management, root cause analysis, and post-incident reviews
  • Excellent problem-solving and analytical skills, with a keen attention to detail.
  • Excellent communication, collaboration, and leadership skills.
  • Seasoned ability to optimize system performance, scalability, and reliability. experience with performance monitoring and tuning tools (for example, Prometheus, Grafana, or New Relic) to identify bottlenecks, analyze performance data, and implement optimization strategies.
  • Seasoned understanding of security principles, best practices, and compliance requirements. experience in designing and implementing security controls, performing security assessments, and ensuring compliance with industry standards.
Academic qualifications and certifications
  • Bachelor's degree or equivalent in Computer Science, Information Technology, or a related field.
  • Relevant certifications, such as AWS Certified DevOps Engineer - Professional, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator (CKA) preferred.
Required experience
  • Seasoned hands-on experience in a Site Reliability Engineering role or related roles, including experience in designing and maintaining highly available and scalable systems.
  • Seasoned hands-on experience with Linux/Unix systems, networking, and system administration is crucial. In-depth knowledge of cloud platforms (such as AWS, Azure, or Google Cloud) and associated services is essential.
  • Seasoned proficiency in multiple programming languages like Python, Java, Go, or Ruby is important for developing and maintaining automation tools, frameworks, and complex system integrations. Expertise in scripting languages like Bash or PowerShell is beneficial.
  • Seasoned understanding of complex infrastructure architectures, including scalable and fault-tolerant designs. experience with infrastructure-as-code tools (such as Terraform or CloudFormation) and containerization technologies (such as Docker or Kubernetes) is essential.
  • Seasoned experience in designing and implementing robust automation frameworks, CI/CD pipelines, and deployment strategies. Proficiency in tools like Jenkins, GitLab CI/CD, or CircleCI to build, test, and deploy applications with a focus on reliability and scalability.
  • Seasoned experience in incident management, troubleshooting complex system issues, and conducting post-incident analysis. Advanced ability to lead incident response efforts, drive root cause analysis, and implement preventive measures.
  • Seasoned understanding of DevOps principles, Agile methodologies, and a strong commitment to continuous improvement and learning. experience in promoting a DevOps culture and driving the adoption of best practices
Workplace type

On-site Working

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

JobCubby • Hyderabad

On-site
INR 1,800,000 - 3,200,000
Senior Associate Site Reliability Engineer
Senior Associate Site Reliability Engineer

NTT DATA, Inc. • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

NTT DATA, Inc. • Hyderabad

On-site
INR 2,500,000 - 4,500,000
Senior Associate Site Reliability Engineer
Senior Associate Site Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Hyderabad

On-site
INR 1,400,000 - 2,000,000
Site Reliability Engineer (SRE) – Core IT Infrastructure
Site Reliability Engineer (SRE) – Core IT Infrastructure

TECEZE • Chennai District

On-site
INR 1,000,000 - 2,000,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Lyzr AI • Bengaluru

Hybrid
INR 1,000,000 - 2,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Resilience and Reliability Engineer
Resilience and Reliability Engineer

EY • Pune District, Gurugram District, Bengaluru

Hybrid
INR 1,800,000 - 2,800,000
SRE Reliability Enginner
SRE Reliability Enginner

NTT DATA North America • Bengaluru Urban

Hybrid
INR 1,200,000 - 2,100,000