SRE Reliability Engineer

NTT DATA BUSINESS SOLUTIONS

Bengaluru

On-site

INR 2,500,000 - 4,000,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NTT DATA BUSINESS SOLUTIONS in Bengaluru, Karnataka, India seeks an experienced Site Reliability Engineer to ensure the availability, reliability, and performance of production systems. The role emphasizes Kubernetes-based operations, observability with Datadog/Prometheus, and Java applications with strong SQL troubleshooting.

You will lead L2/L3 production support, incidents, and automation initiatives across multi-layer stacks, collaborating with Dev, Infra, and Security teams.

Qualifications

  • 5+ years of IT experience in SRE/DevOps roles.
  • Hands-on with Kubernetes and containerized apps.
  • Experience with Datadog or Prometheus for monitoring.
  • Production experience with Java/J2EE microservices.
  • Good JVM troubleshooting and GC tuning.
  • Strong SQL skills for DB issues.
  • Experience managing P1/P2 production incidents.
  • Familiar with REST APIs and distributed systems.
  • Proficient in Linux/Unix and shell scripting.
  • Understanding of CI/CD and deployment practices.
  • Strong analytical and troubleshooting skills.

Responsibilities

  • Own reliability, availability, and operational health of production apps.
  • Provide L2/L3 production support including incident triage and resolution.
  • Monitor apps on Kubernetes including pods, deployments, services, and ingress.
  • Maintain monitoring with Datadog, Prometheus, dashboards, logs, alerts.
  • Define and monitor SLIs, SLOs, SLAs, error budgets.
  • Troubleshoot production issues across Java apps, APIs, microservices, Kubernetes, DBs.
  • Analyze Java logs, JVM performance, memory, threads, and GC issues.
  • Use SQL to investigate production incidents and data issues.
  • Participate in incident management, RCA, post-incident reviews.
  • Identify recurring issues and automate remediation.
  • Build dashboards, alerting, runbooks to reduce MTTR.
  • Drive improvements in MTTR, availability, performance.
  • Automate repetitive ops tasks via scripting/DevOps tooling.
  • Support releases, deployments, rollbacks, validation.
  • Collaborate with Dev, Infra, DevOps, DB, Security, Business teams.
  • Participate in on-call rotations.

Skills

Kubernetes
Datadog
Prometheus
Java/microservices
JVM troubleshooting
SQL
P1/P2 incidents
REST APIs
Linux/Unix
CI/CD
Release management
On-call rotations

Tools

Docker
Helm
GitLab/GitHub Actions
ELK/OpenSearch or Splunk
Ansible/Terraform

Job description

Job Summary

We are looking for an experienced Site Reliability Engineer (SRE) with a strong background in Kubernetes, production support, observability, SQL, and Java-based applications. The role will focus on ensuring the availability, reliability, scalability, and performance of business-critical production systems.

We are currently seeking a SRE Reliability Engineer to join our team in Bangalore, Karnataka, India.

The ideal candidate will combine strong application troubleshooting skills with SRE and DevOps practices, using Datadog and/or Prometheus for monitoring and observability and Kubernetes for managing containerized workloads. A strong understanding of Java applications and relational databases/SQL is essential for diagnosing issues across application, infrastructure, and data layers.

Key Responsibilities
  • Own the reliability, availability, and operational health of business-critical production applications and services.
  • Provide L2/L3 production support, including incident triage, troubleshooting, resolution, and stakeholder communication.
  • Monitor and support applications deployed on Kubernetes, including pods, deployments, services, ingress, resource utilization, scaling, and cluster-related issues.
  • Implement and maintain application and infrastructure monitoring using Datadog, Prometheus, dashboards, metrics, logs, and alerts.
  • Define and monitor SLIs, SLOs, SLAs, error budgets, and service-health indicators for critical applications.
  • Troubleshoot production issues across Java applications, APIs, microservices, Kubernetes, databases, and infrastructure.
  • Analyze Java application logs, exceptions, JVM performance, memory utilization, thread behavior, and garbage collection to identify performance and reliability issues.
  • Use SQL to investigate production incidents, validate data, identify data-related issues, and perform application-level troubleshooting.
  • Participate in incident management, major incident calls, root-cause analysis (RCA), and post-incident reviews.
  • Identify recurring production issues and drive permanent remediation through automation and engineering improvements.
  • Build and enhance monitoring dashboards, alerting mechanisms, and operational runbooks to improve early detection and reduce recovery time.
  • Drive improvements in MTTR, availability, performance, capacity, and production stability.
  • Automate repetitive operational activities using scripting and appropriate DevOps/SRE tooling.
  • Support application releases, production deployments, rollback activities, and post-deployment validation.
  • Work closely with Development, Infrastructure, DevOps, Database, Security, and Business teams to ensure production readiness.
  • Participate in on-call and production support rotations as required.
Required Skills & Experience
  • 5+ years of overall IT experience, with significant experience in SRE, Production Support, Application Support, or DevOps roles.
  • Strong hands-on experience with Kubernetes and containerized applications.
  • Experience with Datadog and/or Prometheus for monitoring, alerting, metrics, and observability.
  • Strong experience supporting Java/J2EE or Java-based microservices applications in production.
  • Good understanding of JVM troubleshooting, application logs, memory, threads, garbage collection, and performance issues.
  • Strong SQL skills with experience troubleshooting relational databases and application data issues.
  • Experience managing P1/P2 production incidents, including incident coordination, RCA, and problem management.
  • Good understanding of REST APIs, microservices, distributed systems, and application integration patterns.
  • Experience with Linux/Unix environments and shell scripting.
  • Understanding of CI/CD pipelines, release management, and deployment practices.
  • Strong analytical and troubleshooting skills with the ability to diagnose issues across multiple technology layers.
Preferred Skills
  • Experience with cloud platforms such as AWS/ Azure.
  • Familiarity with Docker, Helm, GitLab/GitHub Actions, or similar DevOps tooling.
  • Experience with centralized logging platforms such as ELK/OpenSearch or Splunk.
  • Exposure to Infrastructure as Code tools such as Ansible/Terraform.
  • Understanding of load balancing, networking, DNS, certificates, and application security.
  • Experience implementing automation to reduce manual operational effort and production toil.
  • Familiarity with ITIL processes including Incident, Problem, and Change Management.
Key SRE Competencies
  • Production Reliability: Ability to maintain highly available and resilient production services.
  • Observability: Strong understanding of metrics, logs, traces, dashboards, and actionable alerting.
  • Incident Management: Ability to rapidly diagnose and restore services during critical incidents.
  • Problem Management: Strong RCA skills with focus on permanent remediation rather than repeated tactical fixes.
  • Automation: Ability to identify and automate repetitive production-support activities.
  • Performance Engineering: Ability to identify bottlenecks across Java applications, Kubernetes, and databases.
  • Stakeholder Management: Ability to communicate clearly during incidents and work effectively across engineering and business teams.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Senior SRE Technical Specialist
Senior SRE Technical Specialist

United States Digital Space LLC • Karnataka

On-site
INR 1,500,000 - 2,000,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • India

On-site
INR 2,000,000 - 4,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Infobell It Solutions • Hyderabad

On-site
INR 2,800,000 - 4,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
SRE - Site Reliability Engineering
SRE - Site Reliability Engineering

Build & Hire • Pune District

On-site
INR 1,500,000 - 2,300,000
SRE- Production Support
SRE- Production Support

Cloudxtreme • Hyderabad

On-site
INR 2,400,000 - 3,600,000
SRE - AWS DevOPS Engineer
SRE - AWS DevOPS Engineer

Prowess Publishing • Hyderabad

On-site
INR 900,000 - 1,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000