Technology Consultant - Site Reliability Engineer (SRE)

Creative Solutions Services, LLC

Atlanta (GA)

On-site

USD 120,000 - 160,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NTT DATA’s client is seeking an experienced Technology Consultant – Site Reliability Engineer (SRE) to enhance production reliability for enterprise applications. You will work with application engineering, DevOps, cloud, and infrastructure teams to boost availability, scalability, performance, and operational efficiency.

Responsibilities include troubleshooting Kubernetes deployments, implementing observability solutions, supporting Java/Spring Boot microservices, and driving automation to

Qualifications

  • 6+ years of experience in Site Reliability Engineering, DevOps, or Production Engineering/Support.
  • 4+ years of hands-on experience with Kubernetes, Docker, and containerized application environments.
  • 4+ years of experience with Java, Spring Boot, Microservices, and REST APIs.
  • 3+ years of experience with observability and monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, or ELK.

Responsibilities

  • Manage and support business-critical applications running on Kubernetes and containerized platforms.
  • Monitor application and platform health and proactively identify reliability, availability, and performance issues.
  • Troubleshoot Kubernetes deployments, pods, services, networking, configurations, and application issues.
  • Implement and enhance observability solutions covering metrics, logs, traces, dashboards, and alerting.
  • Support and troubleshoot Java/Spring Boot and Microservices-based applications.
  • Perform root cause analysis (RCA) for critical production incidents and implement permanent corrective actions.
  • Define and monitor SLIs, SLOs, SLAs, Error Budgets, and other reliability metrics.
  • Automate repetitive operational activities and identify opportunities to reduce operational TOIL.
  • Participate in incident, problem, change, and production release management activities.
  • Collaborate with engineering teams to improve application resilience, performance, scalability, and fault tolerance.
  • Support CI/CD pipelines and improve application deployment and release processes.
  • Participate in capacity planning, performance tuning, disaster recovery, and production readiness reviews.
  • Develop and maintain operational runbooks, troubleshooting procedures, and technical documentation.

Skills

Site Reliability Engineering
DevOps
Automation
Production Engineering

Tools

Kubernetes
Docker
Java
Spring Boot
Microservices
REST APIs
Splunk
Dynatrace
Prometheus
Grafana
Datadog
ELK
Helm
AWS
Azure
GCP
Linux
Shell scripting
Kafka
Terraform
Ansible
Jenkins
GitHub Actions
GitLab CI
Azure DevOps

Job description

NTT DATA's Client is currently seeking an experienced Technology Consultant – Site Reliability Engineer (SRE) with strong hands-on expertise in Kubernetes, Observability, Java, and production reliability. The ideal candidate will have experience supporting highly available and distributed enterprise applications, troubleshooting complex production issues, and driving automation and reliability improvements.

The role requires close collaboration with application engineering, DevOps, cloud, infrastructure, and support teams to improve application availability, scalability, performance, and operational efficiency.

Day to Day Job Duties
  • Manage and support business-critical applications running on Kubernetes and containerized platforms.
  • Monitor application and platform health and proactively identify reliability, availability, and performance issues.
  • Troubleshoot Kubernetes deployments, pods, services, networking, configurations, and application issues.
  • Implement and enhance observability solutions covering metrics, logs, traces, dashboards, and alerting.
  • Support and troubleshoot Java/Spring Boot and Microservices-based applications.
  • Perform root cause analysis (RCA) for critical production incidents and implement permanent corrective actions.
  • Define and monitor SLIs, SLOs, SLAs, Error Budgets, and other reliability metrics.
  • Automate repetitive operational activities and identify opportunities to reduce operational TOIL.
  • Participate in incident, problem, change, and production release management activities.
  • Collaborate with engineering teams to improve application resilience, performance, scalability, and fault tolerance.
  • Support CI/CD pipelines and improve application deployment and release processes.
  • Participate in capacity planning, performance tuning, disaster recovery, and production readiness reviews.
  • Develop and maintain operational runbooks, troubleshooting procedures, and technical documentation.
Basic Qualifications
  • 6+ years of experience in Site Reliability Engineering, DevOps, or Production Engineering/Support.
  • 4+ years of hands-on experience with Kubernetes, Docker, and containerized application environments.
  • 4+ years of experience with Java, Spring Boot, Microservices, and REST APIs.
  • 3+ years of experience with observability and monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, or ELK.
Nice to Have
  • Strong understanding of SLI, SLO, SLA, Error Budgeting, and SRE principles.
  • Experience with Kubernetes deployment and troubleshooting tools such as Helm.
  • Experience with AWS, Azure, or Google Cloud Platform.
  • Knowledge of Linux/Unix and Shell scripting.
  • Experience with Kafka, IBM MQ, or other messaging technologies.
  • Knowledge of Terraform, Ansible, or other Infrastructure as Code tools.
  • Experience with Jenkins, GitLab CI, GitHub Actions, or Azure DevOps.
  • Experience implementing distributed tracing and application performance monitoring.
  • Knowledge of incident management and ITIL processes.
  • Experience supporting high-volume, highly available, distributed enterprise applications.
  • Strong analytical, troubleshooting, communication, and problem-solving skills.
#LI-NorthAmerica
About NTT DATA:

NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start‑up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com

NTT DATA endeavors to make https://us.nttdata.com accessible to any and all users. If you would like to contact us regarding the accessibility of our website or need assistance completing the application process, please contact us at https://us.nttdata.com/en/contact-us. This contact information is for accommodation requests only and cannot be used to inquire about the status of applications.

NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. For our EEO Policy Statement, please click here. If you'd like more information on your EEO rights under the law, please click here. For Pay Transparency information, please click here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering (SRE) - Maryland, US
Site Reliability Engineering (SRE) - Maryland, US

NTT DATA North America • Baltimore (MD)

On-site
USD 88,000 - 110,000
Medical, dental, vision insurance
401k with company match
Paid time off
Site Reliability Engineering (SRE) - Maryland, US
Site Reliability Engineering (SRE) - Maryland, US

NTT DATA, Inc. • Baltimore (MD)

On-site
USD 88,000 - 110,000
Senior SRE: Kubernetes, Java & Observability
Senior SRE: Kubernetes, Java & Observability

Creative Solutions Services, LLC • Atlanta (GA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer (Onsite Hybrid)
Site Reliability Engineer (Onsite Hybrid)

NTT DATA, Inc. • Plano (TX)

On-site
USD 96,800 - 145,200
Medical, dental, and vision insurance
Flexible spending or health savings account
401(k) program with company match
Lead Infrastructure Engineer
Lead Infrastructure Engineer

Creative Solutions Services, LLC • Irving (TX)

On-site
USD 186,252,000 - 243,560,000
Lead OCP & AWS Cloud Engineer
Lead OCP & AWS Cloud Engineer

Creative Solutions Services, LLC • Irving (TX)

Hybrid
USD 83,000 - 98,000
DevOps Engineer - Westlake, TX
DevOps Engineer - Westlake, TX

NTT DATA Americas • Westlake (LA)

On-site
USD 107,000 - 142,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
+4
DevOps Engineer - Jersey City, NJ
DevOps Engineer - Jersey City, NJ

NTT DATA Americas • Jersey (VA)

On-site
USD 128,000 - 171,000
DevOps Engineer - Plano, TX
DevOps Engineer - Plano, TX

NTT DATA North America • Plano (TX)

On-site
USD 103,000 - 137,000
Software Development Specialist
Software Development Specialist

Creative Solutions Services, LLC • Jersey City (NJ)

On-site
USD 69,000 - 83,000