Technology Consultant - Site Reliability Engineer (SRE)

NTT DATA North America

Atlanta (GA)

On-site

USD 140,000 - 180,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NTT DATA North America is seeking an experienced Technology Consultant - Site Reliability Engineer (SRE) to enhance reliability of distributed enterprise applications. You will collaborate across engineering, DevOps, cloud, and support teams to improve availability, scalability, and efficiency.

Responsibilities include managing Kubernetes-based apps, implementing observability, troubleshooting production issues, and driving automation to reduce toil.

Qualifications

  • 6+ years of experience in Site Reliability Engineering, DevOps, or Production Engineering/Support.
  • 4+ years of hands-on experience with Kubernetes, Docker, and containerized applications.
  • 4+ years of experience with Java, Spring Boot, Microservices, and REST APIs.
  • 3+ years of experience with observability and monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, or ELK.

Responsibilities

  • Manage and support business-critical applications running on Kubernetes and containerized platforms.
  • Monitor application and platform health and proactively identify reliability, availability, and performance issues.
  • Troubleshoot Kubernetes deployments, pods, services, networking, configurations, and application issues.
  • Implement and enhance observability solutions covering metrics, logs, traces, dashboards, and alerting.
  • Support and troubleshoot Java/Spring Boot and Microservices-based applications.
  • Perform root cause analysis (RCA) for critical production incidents and implement permanent corrective actions.
  • Define and monitor SLIs, SLOs, SLAs, Error Budgets, and other reliability metrics.
  • Automate repetitive operational activities and identify opportunities to reduce operational TOIL.
  • Participate in incident, problem, change, and production release management activities.
  • Collaborate with engineering teams to improve application resilience, performance, scalability, and fault tolerance.
  • Support CI/CD pipelines and improve application deployment and release processes.
  • Participate in capacity planning, performance tuning, disaster recovery, and production readiness reviews.
  • Develop and maintain operational runbooks, troubleshooting procedures, and technical documentation.

Skills

Site Reliability
DevOps
Production Engineering
Problem solving

Tools

Kubernetes
Docker
Prometheus
Grafana
Splunk
Dynatrace
Datadog
ELK

Job description

NTT DATA's Client is currently seeking an experienced Technology Consultant - Site Reliability Engineer (SRE) with strong hands-on expertise in Kubernetes, Observability, Java, and production reliability. The ideal candidate will have experience supporting highly available and distributed enterprise applications, troubleshooting complex production issues, and driving automation and reliability improvements.

The role requires close collaboration with application engineering, DevOps, cloud, infrastructure, and support teams to improve application availability, scalability, performance, and operational efficiency.

Day to Day Job Duties
  • Manage and support business-critical applications running on Kubernetes and containerized platforms.
  • Monitor application and platform health and proactively identify reliability, availability, and performance issues.
  • Troubleshoot Kubernetes deployments, pods, services, networking, configurations, and application issues.
  • Implement and enhance observability solutions covering metrics, logs, traces, dashboards, and alerting.
  • Support and troubleshoot Java/Spring Boot and Microservices-based applications.
  • Perform root cause analysis (RCA) for critical production incidents and implement permanent corrective actions.
  • Define and monitor SLIs, SLOs, SLAs, Error Budgets, and other reliability metrics.
  • Automate repetitive operational activities and identify opportunities to reduce operational TOIL.
  • Participate in incident, problem, change, and production release management activities.
  • Collaborate with engineering teams to improve application resilience, performance, scalability, and fault tolerance.
  • Support CI/CD pipelines and improve application deployment and release processes.
  • Participate in capacity planning, performance tuning, disaster recovery, and production readiness reviews.
  • Develop and maintain operational runbooks, troubleshooting procedures, and technical documentation.
Basic Qualifications
  • 6+ years of experience in Site Reliability Engineering, DevOps, or Production Engineering/Support.
  • 4+ years of hands-on experience with Kubernetes, Docker, and containerized application environments.
  • 4+ years of experience with Java, Spring Boot, Microservices, and REST APIs.
  • 3+ years of experience with observability and monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, or ELK.
Nice to Have
  • Strong understanding of SLI, SLO, SLA, Error Budgeting, and SRE principles.
  • Experience with Kubernetes deployment and troubleshooting tools such as Helm.
  • Experience with AWS, Azure, or Google Cloud Platform.
  • Knowledge of Linux/Unix and Shell scripting.
  • Experience with Kafka, IBM MQ, or other messaging technologies.
  • Knowledge of Terraform, Ansible, or other Infrastructure as Code tools.
  • Experience with Jenkins, GitLab CI, GitHub Actions, or Azure DevOps.
  • Experience implementing distributed tracing and application performance monitoring.
  • Knowledge of incident management and ITIL processes.
  • Experience supporting high-volume, highly available, distributed enterprise applications.
  • Strong analytical, troubleshooting, communication, and problem-solving skills.
About NTT DATA

NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com

NTT DATA endeavors to make https://us.nttdata.com accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact us at https://us.nttdata.com/en/contact-us. This contact information is for accommodation requests only and cannot be used to inquire about the status of applications. NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. For our EEO Policy Statement, please click here. If you'd like more information on your EEO rights under the law, please click here. For Pay Transparency information, please click here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technology Consultant - Site Reliability Engineer (SRE)
Technology Consultant - Site Reliability Engineer (SRE)

Creative Solutions Services, LLC • Atlanta (GA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer- REMOTE - Onsite Training
Site Reliability Engineer- REMOTE - Onsite Training

NTT DATA North America • Memphis (TN)

On-site
USD 87,000 - 151,000
Health insurance
Dental and vision insurance
401k with company match
+5
Technology Consultant - Cloud & Full Stack Engineer
Technology Consultant - Cloud & Full Stack Engineer

NTT DATA North America • Atlanta (GA)

On-site
USD 76,000 - 85,000
Technology Consultant - Cloud & Full Stack Engineer
Technology Consultant - Cloud & Full Stack Engineer

Creative Solutions Services, LLC • Atlanta (GA)

On-site
USD 76,000 - 85,000
Medical, dental, and vision insurance
HSA/FSA options
401k program with company match
+2
Site Reliability Engineer (Onsite Hybrid)
Site Reliability Engineer (Onsite Hybrid)

NTT DATA, Inc. • Plano (TX)

On-site
USD 96,800 - 145,200
Medical, dental, and vision insurance
Flexible spending or health savings account
401(k) program with company match
Senior SRE Consultant: Kubernetes & Observability
Senior SRE Consultant: Kubernetes & Observability

NTT DATA North America • Atlanta (GA)

On-site
USD 140,000 - 180,000
Software Development Specialist
Software Development Specialist

NTT DATA North America • Jersey City (NJ)

Hybrid
USD 69,000 - 83,000
Java Data Engineer
Java Data Engineer

Creative Solutions Services, LLC • Nashville (TN)

On-site
USD 69,000 - 90,000
Medical insurance
Dental insurance
Vision insurance
+2
Tech Consultant - Lead Full-Stack Engineer
Tech Consultant - Lead Full-Stack Engineer

NTT DATA North America • New York (NY)

Hybrid
USD 118,000 - 157,000
Medical, dental, and vision insurance
Employer 401k with company match
Paid time off
+3
Lead OCP & AWS Cloud Engineer
Lead OCP & AWS Cloud Engineer

Creative Solutions Services, LLC • Irving (TX)

Hybrid
USD 83,000 - 98,000