Site Reliability Engineer

NTT Limited

Town of Texas (WI)

Remote

USD 80,000 - 148,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NTT DATA is seeking a seasoned Site Reliability Engineer to ensure reliability, availability, and performance of our systems and infrastructure. You will collaborate with development and operations teams to automate processes and drive resilience across environments.

The role emphasizes incident management, automation tooling, and CI/CD practices, with a remote work option for US-based candidates. Candidates should have strong Linux/Unix, scripting, and cloud experience.

Qualifications

  • Bachelor’s degree in CS/IT or related field.
  • Certifications in cloud or DevOps are a plus (AWS, Google Cloud, or Kubernetes).
  • 5+ years in Site Reliability Engineering or related roles.
  • Strong Linux/Unix administration and networking knowledge.
  • Proficiency in Python, Go, Java, or Ruby for automation.
  • Experience with IaC (Terraform, CloudFormation) and containers (Docker, Kubernetes).

Responsibilities

  • Monitors system health, performance metrics, and alerts to identify incidents and restore services.
  • Implements incident response processes to minimize downtime.
  • Designs, develops, and maintains automation tools and scripts.
  • Implements infrastructure‑as‑code for consistency and repeatability.
  • Optimizes resources and configurations for better performance and scalability.
  • Uses Prometheus/Grafana/New Relic for monitoring and tuning.
  • Leads incident response and performs root-cause analysis to prevent recurrences.
  • Collaborates with teams to ensure reliability goals and security standards.

Skills

Linux/Unix
Networking
SRE fundamentals
Python
Go
Java
Cloud platforms
Terraform
Docker
Kubernetes
Prometheus
Grafana
New Relic
Jenkins
Post-incident analysis

Education

Bachelor's degree in Computer Science, IT, or related field

Tools

Terraform
Docker
Kubernetes
Jenkins
GitLab CI/CD
Prometheus
Grafana
New Relic

Job description

Make an impact with NTT DATA Join a company that is pushing the boundaries of what is possible. We are renowned for our technical excellence and leading innovations, and for making a difference to our clients and society. Our workplace embraces diversity and inclusion – it’s a place where you can grow, belong and thrive.

Your day at NTT DATA

The Site Reliability Engineer (SRE) is a seasoned subject matter expert, responsible for ensuring the reliability, availability, and performance of company systems and infrastructure. This Site Reliability Engineer (SRE) works closely with development teams, operations teams, and other stakeholders to enhance system resiliency, automate processes, and improve overall system reliability.

Key responsibilities
  • Monitors system health, performance metrics, and alerts to identify and respond to incidents promptly and diagnoses issues, troubleshoots problems, and restores services in a timely manner.
  • Implements incident response processes to minimize downtime and improve system availability.
  • Designs, develops, and maintains automation tools, scripts, and processes to streamline system management tasks, deployments, and configuration changes.
  • Implements infrastructure-as-code principles to ensure consistency and repeatability.
  • Optimizes system resources, configurations, and processes to enhance performance, scalability, and efficiency.
  • Uses monitoring tools and performance testing to identify bottlenecks and implement optimizations.
  • Collaborates with teams to forecast system resource needs, plans for capacity growth, and ensures adequate scalability.
  • Leads incident response efforts, coordinates with cross-functional teams, and drives the resolution of system issues.
  • Performs thorough post-incident analysis to identify root causes and implements preventive measures to minimize future incidents.
  • Identifies opportunities for automation and drives the implementation of self-healing, monitoring, and deployment of automation tools and frameworks.
  • Continuously improves operational efficiency, system reliability, and availability through process enhancements and automation.
  • Ensures consistency across environments, tracks changes, and enforces configuration standards.
  • Works closely with development teams, operations teams, and other stakeholders to ensure effective collaboration, knowledge sharing, and alignment on reliability goals.
  • Implements security best practices, works with security teams to assess and address vulnerabilities, and ensures compliance with security standards and regulations.
  • Performs any other related task as required.
To thrive in this role, you need to have:
  • Seasoned technical expertise in Linux/Unix systems, networking, and system administration.
  • Seasoned proficiency in scripting or programming languages, such as Python, Go, Java, or Ruby.
  • Seasoned knowledge of cloud platforms (such as AWS, Azure, or Google Cloud) and associated services.
  • Seasoned proven expertise in performance monitoring, optimization, and troubleshooting using tools such as Prometheus, Grafana, or New Relic.
  • Seasoned expertise in incident management, root cause analysis, and post-incident reviews.
  • Excellent problem-solving and analytical skills, with a keen attention to detail.
  • Excellent communication, collaboration, and leadership skills.
  • Seasoned ability to optimize system performance, scalability, and reliability.
  • experience with performance monitoring and tuning tools (for example, Prometheus, Grafana, or New Relic) to identify bottlenecks, analyze performance data, and implement optimization strategies.
  • Seasoned understanding of security principles, best practices, and compliance requirements.
  • experience in designing and implementing security controls, performing security assessments, and ensuring compliance with industry standards.
Academic qualifications and certifications:
  • Bachelor's degree or equivalent in Computer Science, Information Technology, or a related field.
  • Relevant certifications, such as AWS Certified DevOps Engineer - Professional, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator (CKA) preferred.
Required experience:
  • Seasoned hands‑on experience in a Site Reliability Engineering role or related roles, including experience in designing and maintaining highly available and scalable systems.
  • Seasoned hands‑on experience with Linux/Unix systems, networking, and system administration is crucial.
  • In-depth knowledge of cloud platforms (such as AWS, Azure, or Google Cloud) and associated services is essential.
  • Seasoned proficiency in multiple programming languages like Python, Java, Go, or Ruby is important for developing and maintaining automation tools, frameworks, and complex system integrations.
  • Expertise in scripting languages like Bash or PowerShell is beneficial.
  • Seasoned understanding of complex infrastructure architectures, including scalable and fault‑tolerant designs.
  • experience with infrastructure-as‑code tools (such as Terraform or CloudFormation) and containerization technologies (such as Docker or Kubernetes) is essential.
  • Seasoned experience in designing and implementing robust automation frameworks, CI/CD pipelines, and deployment strategies.
  • Proficiency in tools like Jenkins, GitLab CI/CD, or CircleCI to build, test, and deploy applications with a focus on reliability and scalability.
  • Seasoned experience in incident management, troubleshooting complex system issues, and conducting post-incident analysis.
  • Advanced ability to lead incident response efforts, drive root cause analysis, and implement preventive measures.
  • Seasoned understanding of DevOps principles, Agile methodologies, and a strong commitment to continuous improvement and learning.
  • experience in promoting a DevOps culture and driving the adoption of best practices

NTT DATA provides a reasonable range of compensation for U.S.-based positions. The starting pay range for this remote role is $80,000.00 - 114,000.00 - 148,000.00 USD Annual. This range reflects the minimum and maximum target compensation for the position across all US locations. Actual compensation will depend on a number of factors, including the candidate’s actual work location, relevant experience, technical skills, and other qualifications.

Workplace type: Remote Working

About NTT DATA

NTT DATA is a $30+ billion business and technology services leader, serving 75% of the Fortune Global 100. We are committed to accelerating client success and positively impacting society through responsible innovation. We are one of the world’s leading AI and digital infrastructure providers, with unmatched capabilities in enterprise-scale AI, cloud, security, connectivity, data centers and application services. Our consulting and industry solutions help organizations and society move confidently and sustainably into the digital future. As a Global Top Employer, we have experts in more than 70 countries. We also offer clients access to a robust ecosystem of innovation centers as well as established and start-up partners. NTT DATA is part of NTT Group, which invests over $3 billion each year in R&D.

Equal Opportunity Employer

NTT DATA is proud to be an Equal Opportunity Employer with a global culture that embraces diversity. We are committed to providing an environment free of unfair discrimination and harassment. We do not discriminate based on age, race, colour, gender, sexual orientation, religion, nationality, disability, pregnancy, marital status, veteran status, or any other protected category.

Join our growing global team and accelerate your career with us.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

NTT Limited • Michigan

On-site
USD 118,000 - 218,000
Site Reliability Engineer- REMOTE - Onsite Training
Site Reliability Engineer- REMOTE - Onsite Training

NTT DATA North America • Memphis (TN)

On-site
USD 87,000 - 151,000
Health insurance
Dental and vision insurance
401k with company match
+5
DevOps Engineer
DevOps Engineer

NTT Limited • Town of Texas (WI)

Remote
USD 92,000 - 170,000
Remote working
Manager, Data Centre Operations
Manager, Data Centre Operations

NTT America, Inc. • Austin (TX)

Hybrid
USD 140,000 - 175,000
Medical, dental, vision insurance
401k with company match
Paid time off
Technology Consultant - Site Reliability Engineer (SRE)
Technology Consultant - Site Reliability Engineer (SRE)

NTT DATA North America • Atlanta (GA)

On-site
USD 140,000 - 180,000
Technology Consultant - Site Reliability Engineer (SRE)
Technology Consultant - Site Reliability Engineer (SRE)

JobDiva, Inc. • Atlanta (GA)

On-site
USD 120,000 - 160,000
Senior Java Spring Boot Developer (FTE / Hybrid)
Senior Java Spring Boot Developer (FTE / Hybrid)

NTT America, Inc. • Town of Charlotte (NY)

Hybrid
USD 97,000 - 145,000
Director, Program Manager - Remote in US
Director, Program Manager - Remote in US

NTT America, Inc. • Dallas (TX)

Remote
USD 117,000 - 208,000
Medical, dental, and vision insurance
401k with company match
Paid time off
Senior SRE: Kubernetes, Java & Observability
Senior SRE: Kubernetes, Java & Observability

NTT DATA • Atlanta (GA)

On-site
USD 120,000 - 160,000
Technology Sales Specialist, Support Services
Technology Sales Specialist, Support Services

Urban Ridge Supplies • New York (NY)

Remote
USD 115,000 - 213,000