SRE Reliability Engineer

NTT DATA, Inc.

Bengaluru

Hybrid

INR 2,500,000 - 4,000,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NTT DATA is seeking an experienced Site Reliability Engineer (SRE) to ensure the availability, reliability, and performance of business-critical production systems. You will work with Kubernetes, Java applications, and relational databases, using Datadog/Prometheus for observability and incident response to drive automation and reliability improvements.

The role emphasizes L2/L3 production support, incident management, and collaboration with DevOps, Infra, and Security teams.

Qualifications

  • 5+ years of IT experience in SRE/DevOps roles.
  • Strong Kubernetes and containerized app experience.
  • Experience with Datadog/Prometheus for monitoring and alerting.
  • Production support for Java/J2EE or microservices in production.

Responsibilities

  • Own reliability and operational health of production apps and services.
  • Provide L2/L3 production support including incident triage.
  • Monitor apps on Kubernetes; manage pods, deployments, services, and scaling.
  • Implement and maintain monitoring dashboards, logs, and alerts.
  • Define SLIs, SLOs, SLAs and recovery objectives.
  • Troubleshoot across Java apps, APIs, microservices, and databases.
  • Analyze JVM, memory, GC, and thread behavior for performance issues.
  • Use SQL for production data investigations and troubleshooting.
  • Participate in incident management and RCA processes.
  • Automate repeated operations and support releases and deployments.

Skills

Kubernetes
Datadog
Prometheus
Java/J2EE
SQL
Linux/Unix
CI/CD
Incident Management
Observability
REST APIs

Tools

Datadog
Prometheus

Job description

Site Reliability Engineer (SRE) – Kubernetes / Observability / Java
Role Summary

We are looking for an experienced Site Reliability Engineer (SRE) with a strong background in Kubernetes, production support, observability, SQL, and Java-based applications. The role will focus on ensuring the availability, reliability, scalability, and performance of business-critical production systems.

The ideal candidate will combine strong application troubleshooting skills with SRE and DevOps practices, using Datadog and/or Prometheus for monitoring and observability and Kubernetes for managing containerized workloads. A strong understanding of Java applications and relational databases/SQL is essential for diagnosing issues across application, infrastructure, and data layers.

Key Responsibilities
  • Own the reliability, availability, and operational health of business-critical production applications and services.
  • Provide L2/L3 production support, including incident triage, troubleshooting, resolution, and stakeholder communication.
  • Monitor and support applications deployed on Kubernetes, including pods, deployments, services, ingress, resource utilization, scaling, and cluster-related issues.
  • Implement and maintain application and infrastructure monitoring using Datadog, Prometheus, dashboards, metrics, logs, and alerts.
  • Define and monitor SLIs, SLOs, SLAs, error budgets, and service-health indicators for critical applications.
  • Troubleshoot production issues across Java applications, APIs, microservices, Kubernetes, databases, and infrastructure.
  • Analyze Java application logs, exceptions, JVM performance, memory utilization, thread behavior, and garbage collection to identify performance and reliability issues.
  • Use SQL to investigate production incidents, validate data, identify data-related issues, and perform application-level troubleshooting.
  • Participate in incident management, major incident calls, root-cause analysis (RCA), and post-incident reviews.
  • Identify recurring production issues and drive permanent remediation through automation and engineering improvements.
  • Build and enhance monitoring dashboards, alerting mechanisms, and operational runbooks to improve early detection and reduce recovery time.
  • Drive improvements in MTTR, availability, performance, capacity, and production stability.
  • Automate repetitive operational activities using scripting and appropriate DevOps/SRE tooling.
  • Support application releases, production deployments, rollback activities, and post-deployment validation.
  • Work closely with Development, Infrastructure, DevOps, Database, Security, and Business teams to ensure production readiness.
  • Participate in on-call and production support rotations as required.
Required Skills & Experience
  • 5+ years of overall IT experience , with significant experience in SRE, Production Support, Application Support, or DevOps roles.
  • Strong hands-on experience with Kubernetes and containerized applications.
  • Experience with Datadog and/or Prometheus for monitoring, alerting, metrics, and observability.
  • Strong experience supporting Java/J2EE or Java-based microservices applications in production.
  • Good understanding of JVM troubleshooting, application logs, memory, threads, garbage collection, and performance issues.
  • Strong SQL skills with experience troubleshooting relational databases and application data issues.
  • Experience managing P1/P2 production incidents, including incident coordination, RCA, and problem management.
  • Good understanding of REST APIs, microservices, distributed systems, and application integration patterns.
  • Experience with Linux/Unix environments and shell scripting.
  • Understanding of CI/CD pipelines, release management, and deployment practices.
  • Strong analytical and troubleshooting skills with the ability to diagnose issues across multiple technology layers.
Preferred Skills
  • Experience with cloud platforms such as AWS/ Azure.
  • Familiarity with Docker, Helm, GitLab/GitHub Actions, or similar DevOps tooling.
  • Experience with centralized logging platforms such as ELK/OpenSearch or Splunk.
  • Exposure to Infrastructure as Code tools such as Ansible/Terraform.
  • Understanding of load balancing, networking, DNS, certificates, and application security.
  • Experience implementing automation to reduce manual operational effort and production toil.
  • Familiarity with ITIL processes including Incident, Problem, and Change Management.
  • Production Reliability: Ability to maintain highly available and resilient production services.
  • Observability: Strong understanding of metrics, logs, traces, dashboards, and actionable alerting.
  • Incident Management: Ability to rapidly diagnose and restore services during critical incidents.
  • Problem Management: Strong RCA skills with focus on permanent remediation rather than repeated tactical fixes.
  • Automation: Ability to identify and automate repetitive production-support activities.
  • Performance Engineering: Ability to identify bottlenecks across Java applications, Kubernetes, and databases.
  • Stakeholder Management: Ability to communicate clearly during incidents and work effectively across engineering and business teams.
Success Measures

Success in this role will be measured through improvements in system availability, SLO adherence, MTTR, incident recurrence, application performance, alert quality, production stability, and reduction of manual operational effort.

About NTT DATA

NTT DATA is a $30 billion business and technology services leader, serving 75% of the Fortune Global 100. We are committed to accelerating client success and positively impacting society through responsible innovation. We are one of the world's leading AI and digital infrastructure providers, with unmatched capabilities in enterprise-scale AI, cloud, security, connectivity, data centers and application services. our consulting and Industry solutions help organizations and society move confidently and sustainably into the digital future. As a Global Top Employer, we have experts in more than 50 countries. We also offer clients access to a robust ecosystem of innovation centers as well as established and start-up partners.NTT DATA is a part of NTT Group, which invests over $3 billion each year in R&D.

Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or hybrid work options, these arrangements are subject to change based on client requirements. For employees near an NTT DATA office or client site, in-office attendance may be required for meetings or events, depending on business needs. At NTT DATA, we are committed to staying flexible and meeting the evolving needs of both our clients and employees. NTT DATA recruiters will never ask for payment or banking information and will only use @nttdata.com, @nttdatafed.com and @talent.nttdataservices.com email addresses. If you are requested to provide payment or disclose banking information, please submit a contact us form, https://us.nttdata.com/en/contact-us .

NTT DATA endeavors to make https://us.nttdata.com accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact us at https://us.nttdata.com/en/contact-us . This contact information is for accommodation requests only and cannot be used to inquire about the status of applications. NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. For our EEO Policy Statement, please click here . If you'd like more information on your EEO rights under the law, please click here . For Pay Transparency information, please click here .


Job Segment: Java, Developer, Database, SQL, Linux, Technology

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Reliability Enginner
SRE Reliability Enginner

NTT DATA North America • Bengaluru Urban

Hybrid
INR 1,200,000 - 2,100,000
SRE Reliability Engineer
SRE Reliability Engineer

NTT DATA North America • Bengaluru Urban

On-site
INR 2,500,000 - 4,500,000
Site Reliability Engineer
Site Reliability Engineer

JobCubby • Hyderabad

On-site
INR 1,800,000 - 3,200,000
Site Reliability Engineer
Site Reliability Engineer

NTT DATA, Inc. • Hyderabad

On-site
INR 2,500,000 - 4,500,000
Technical Consultant - Java Full Stack Developer
Technical Consultant - Java Full Stack Developer

NTT DATA North America • Pune District

On-site
INR 1,200,000 - 1,600,000
Senior Engineering Support Specialist - Controls Engineering
Senior Engineering Support Specialist - Controls Engineering

NTT DATA Global Delivery Services Ltd • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Programming Language Knowledge - Java
Programming Language Knowledge - Java

NTT DATA North America • Bengaluru Urban

On-site
INR 600,000 - 1,200,000
Devops Consultant
Devops Consultant

NTT DATA, Inc. • Bengaluru

Hybrid
INR 1,200,000 - 1,600,000
Senior SDET (C# + Selenium)
Senior SDET (C# + Selenium)

NTT DATA, Inc. • Hyderabad

On-site
INR 1,800,000 - 2,500,000
Engineers (Agentic / Automation)
Engineers (Agentic / Automation)

NTT DATA, Inc. • Gurgaon

On-site
INR 3,500,000 - 5,500,000