Senior Site Reliability Engineer

nexocean

Hyderabad

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology solutions company in Hyderabad is seeking a Senior Site Reliability Engineer to design and maintain Kubernetes clusters, automate deployments, and resolve complex incidents. The ideal candidate has extensive experience in Site Reliability Engineering and is proficient in scripting with Go, Bash, and Python. Responsibilities include leading incident responses, optimizing CI/CD pipelines, and ensuring operational excellence across hybrid infrastructures. This role offers a dynamic environment for driving innovation and improving system performance.

Qualifications

  • 5–10 years of experience in Site Reliability Engineering roles.
  • Extensive hands-on experience with Kubernetes.
  • Proficient in Go, Bash, and Python for automation.
  • Experience with CI/CD deployment processes.

Responsibilities

  • Design and maintain large-scale Kubernetes clusters.
  • Automate infrastructure deployment using IaC tools.
  • Lead incident response and root cause analysis efforts.
  • Optimize CI/CD pipelines for seamless software delivery.
  • Enhance system performance through monitoring and analysis.

Skills

Kubernetes expertise
Scripting in Go
Scripting in Python
Infrastructure as Code
Analytical thinking
Collaboration skills

Education

Bachelor’s degree in computer science, Software Engineering, or related field

Tools

Terraform
GitLab
Prometheus
Grafana

Job description

We are excited to announce an opening for Sr Engineer, Site Reliability at TMUS Global Solutions. Please find below the details of the role and its responsibilities.

Job Description
About the Role

The Senior Site Reliability Engineer for Container Platforms plays a crucial role in shaping and maintaining the foundational infrastructure that powers our next-generation platforms and services. They design, implement, and manage large-scale Kubernetes clusters and related automation systems that ensure high availability, scalability, and reliability across technology ecosystem. They also utilize their strong problem‑solving and analytical skills to automate processes, reducing manual effort and preventing operational incidents. Their expertise in Kubernetes and scripting languages, incident response management, and various tech tools contributes to the robustness and efficiency of our systems. By continuously learning new skills and technologies, they adapt to changing circumstances and drive innovation. Their work and expertise contribute significantly to the stability and performance of digital infrastructure. They are also responsible for diagnosing and resolving complex issues across networking, storage, and compute layers, driving continuous improvement through data‑driven insights and DevOps best practices. This engineer is also responsible for contributing to the overall architecture and strategy of technical systems, mentoring junior engineers, and ensuring solutions are aligned with business and technical goals.

What You’ll Do
  • Design, build, and maintain large-scale, production‑grade Kubernetes (K8s) clusters to ensure high availability, scalability, and security across hybrid infrastructure.
  • Develop and manage Infrastructure as Code (IaC) using tools such as Terraform, CloudFormation, and ARM templates, enabling consistent and automated infrastructure deployment across AWS, Azure, and on‑premises data centers.
  • Resolve platform‑related customer tickets by diagnosing and addressing infrastructure, deployment, and performance issues to ensure reliability and seamless user experience.
  • Lead incident response, root cause analysis (RCA), and post‑mortems, implementing automation to prevent recurrence.
  • Implement and optimize CI/CD pipelines leveraging GitLab, Argo, and Flux to support seamless software delivery, continuous integration, and progressive deployment strategies.
  • Automate system operations through scripting and development in Go, Bash, and Python, driving efficiency, repeatability, and reduced operational overhead.
  • Monitor, analyze, and enhance system performance, proactively identifying bottlenecks and ensuring reliability through data‑driven observability and capacity planning.
  • Troubleshoot complex issues across the full stack—network, storage, compute, and application layers using advanced tools like pcap, telnet, and Linux‑native diagnostics.
  • Apply deep Kubernetes expertise to diagnose, resolve, and prevent infrastructure‑related incidents while mentoring team members on container orchestration best practices.
  • Drive a culture of automation, resilience, and continuous improvement, contributing to the evolution of platform engineering and cloud infrastructure strategies.
  • Drive innovation by recommending new technologies, frameworks, and tools.
  • Perform additional duties and strategic projects as assigned.
What You’ll Bring
  • Bachelor’s degree in computer science, Software Engineering, or related field.
  • 5–10 years of hands‑on experience in Site Reliability Engineering roles supporting large‑scale, production‑grade systems.
  • Extensive hands‑on experience with Kubernetes (K8s)—including cluster provisioning, scaling, upgrades, and performance tuning in both on‑premises and multi‑cloud environments.
  • Proficiency in scripting and programming with Go, Bash, or Python to automate operational tasks and develop scalable infrastructure solutions.
  • Hands‑on experience with observability tools (monitoring, alerting, logging, and tracing) to maintain reliability and operational excellence.
  • Solid understanding of CI/CD concepts and practical experience with GitLab pipelines or similar tools for automated deployments and continuous delivery.
  • Strong understanding of cloud architecture and DevOps best practices.
  • Strong analytical thinking and collaborative problem‑solving skills.
  • Excellent communication and documentation abilities.
Must Have Skills
  • Infrastructure as Code (IaC) & Automation using Terraform, Ansible, CloudFormation, ARM Templates.
  • Expert‑level knowledge of Linux administration, performance tuning, and troubleshooting.
  • Strong programming and automation skills in Go, Python, or Bash, coupled with hands‑on experience in CI/CD pipelines.
  • Experience with monitoring and logging tools like Prometheus, Grafana.
Nice To Have
  • Certifications in Kubernetes, cloud platforms, or DevOps practices.
  • Performance Optimization and Capacity Planning.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr Engineer, Site Reliability [T500-27847]
Sr Engineer, Site Reliability [T500-27847]

TMUS Global Solutions • Hyderabad

On-site
INR 2,500,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Headout • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Saama • Chennai District

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Lead
Senior Site Reliability Lead

Generac • Pune District

On-site
INR 3,000,000 - 6,500,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Stryker Group • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer | Global Platform Engineering | Hyderabad
Senior Site Reliability Engineer | Global Platform Engineering | Hyderabad

CareerXperts Consulting • Hyderabad

On-site
INR 1,500,000 - 1,900,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

UST • Pune District

On-site
INR 1,800,000 - 3,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Poshmark • Chennai

On-site
INR 1,200,000 - 1,800,000
Senior DevOps Engineer
Senior DevOps Engineer

Responsive • Bengaluru

On-site
INR 3,500,000 - 7,000,000