Senior Site Reliability Engineer

Nexthink India Digital Experience

Bengaluru

Hybrid

INR 2,600,000 - 4,800,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Nexthink India Digital Experience is seeking a Senior Site Reliability Engineer to strengthen our cloud-native platform, ensuring scalable, reliable, and secure SaaS operations in a multi-tenant environment. You will collaborate with Product and Platform teams to drive observability, automation, and resilient systems.

The role requires 5+ years in SRE/Platform, strong scripting, Terraform, and Kubernetes expertise, with on-call responsibilities and a focus on incident prevention and rapid

Qualifications

  • Bachelor's degree in Computer Science or equivalent practical experience.
  • 5+ years of experience as a Site Reliability Engineer or Platform Engineer.
  • Strong hands-on experience with public cloud services (AWS, GCP, Azure) and supporting SaaS products.
  • Strong programming or scripting skills (e.g., Python, Go, Bash) and experience with infrastructure-as-code (Terraform).
  • Proficiency with Kubernetes, container-based deployment (Docker) and related ecosystems (Helm).
  • Experience supporting multi-tenant microservices architectures.
  • Experience with CI/CD pipelines & tools (Jenkins, GitHub Actions, GitLab CI, FluxCD, Crossplane).
  • Experience with monitoring solutions (Datadog).
  • Comfortable on-call, incident management, and post-incident reviews.
  • Knowledge of Linux, networking, and cloud architectures (VPC, subnets, firewalls, load balancers).
  • Exposure to SOC 2, ISO 27001, HIPAA; FedRAMP a plus.
  • Chaos engineering or resilience testing experience.
  • Excellent communication and collaboration skills; fluent English.

Responsibilities

  • Implement and manage cloud-native systems (AWS) using automation.
  • Operate and enhance Kubernetes clusters, deployment pipelines, and service meshes.
  • Design, build, and maintain infrastructure powering a multi-tenant SaaS platform with reliability and security.
  • Define and maintain SLOs, SLAs, and error budgets; address availability and performance issues.
  • Develop infrastructure-as-code for repeatable provisioning (Terraform).
  • Build internal platform tools and automation for provisioning and monitoring.
  • Monitor infrastructure and applications to ensure high-quality user experiences.
  • Participate in a rotating on-call schedule; respond to incidents and coordinate responses.
  • Act as Incident Commander during on-call duty and coordinate cross-team responses.
  • Improve incident response processes to reduce MTTR/MTTD; diagnose issues and prevent recurrences.
  • Embed observability and fault tolerance into service design with engineers.
  • Automate runbooks, health checks, and alerting for reliable operations.
  • Support automated testing, canary deployments, and rollback strategies for safe releases.
  • Contribute to security and compliance automation and cost optimization.

Skills

Strong problem-solving
English communication
Collaborative mindset
Self-driven
Agile development

Education

Bachelor's degree in Computer Science or equivalent

Tools

Python
Go
Bash
Terraform
Kubernetes
Docker
Helm
Istio
Jenkins
GitHub Actions
GitLab CI
FluxCD
Crossplane
Datadog

Job description

Company Description

Nexthink leads digital employee experience management software. The company provides IT leaders with unprecedented insight, allowing them to see, diagnose and fix issues at scale impacting employees anywhere, with any applicationor network, before employees notice the issue. As the first solutionto help IT move from reactive problem-solving to proactive optimisation, Nexthink enables its more than 1,300 customers to deliver better digital experiences to more than 18 millionemployees. Dual-headquartered in Lausanne, Switzerland and Boston, Massachusetts, Nexthink has 9 offices worldwide.

Job Description

We are looking for an experienced, proactive and innovative professional who is keen to join as a Senior Site Reliability Engineer! The mission of Nexthink's SRE team is to strengthen our infrastructure and enhance our ability to deploy, monitor, and scale systems effectively and reliably. They work closely with over 50 Product Engineering teams that develop our products and services, as well as with the Technical Platform Engineering, Security and Architecture teams to understand the reliability requirements, design and implement solutions, and promote them for adoption and usage.

Join our vibrant team of diverse and experienced engineers where cutting-edge technology meets innovation. Be a part of Nexthink's Digital Employee Experience technological revolution, ensuring our global customers enjoy a seamless user experience.

As a Senior Site Reliability Engineer, you will:
  • Implement and manage cloud-native systems (AWS) using best-in-class tools and automation.
  • Operate and enhance Kubernetes clusters, deployment pipelines, and service meshes to support rapid delivery cycles.
  • Design, build, and maintain the infrastructure powering our multi-tenant SaaS platform with reliability, security, and scalability in mind.
  • Define and maintain SLOs, SLAs, and error budgets, and proactively address availability and performance issues.
  • Develop infrastructure-as-code (Terraform or similar) for repeatable and auditable provisioning.
  • Build internal platform tools and automation to support provisioning, monitoring, and operational efficiency.
  • Monitor infrastructure and applications, ensuring high-quality user experiences.
  • Participate in a shared on-call rotation, responding to incidents, troubleshooting outages, and driving timely resolution and communication.
  • Act as an Incident Commander during on-call duty and coordinate cross-team responses effectively to maintain an SLA.
  • Drive and refine incident response processes, reducing Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR).
  • Diagnose and resolve complex issues independently, minimising the need for external escalation.
  • Work closely with software engineers to embed observability, fault tolerance, and reliability principles into service design.
  • Automate runbooks, health checks, and alerting to support reliable operations with minimal manual intervention.
  • Support automated testing, canary deployments, and rollback strategies to ensure safe, fast, and reliable releases.
  • Contribute to security best practices, compliance automation, and cost optimisation.
Qualifications
  • Minimum Bachelor's degree in Computer Science or equivalent practical experience.
  • 5+ years of experience as a Site Reliability Engineer or Platform Engineer with strong knowledge of software development best practices.
  • Strong hands-on experience with public cloud services (AWS, GCP, Azure) and supporting SaaS products.
  • Strong programming or scripting skills (e.g., Python, Go, Bash...) and experience with infrastructure-as-code (e.g., Terraform).
  • Proficiency with Kubernetes, container-based deployment (e.g., Docker) and related ecosystems (e.g., Helm).
  • Experience supporting multi-tenant microservices architectures.
  • Experience with CI/CD pipelines & tools (e.g., Jenkins, GitHub Actions, GitLab CI, FluxCD, Crossplane).
  • Experience with managing monitoring solutions (e.g. Datadog).
  • Comfortable participating in a rotating on-call schedule, managing critical incidents, and leading post-incident reviews.
  • At ease with operating and managing production systems, striking the right balance between urgency and methodology.
  • Strong system-level troubleshooting skills and a proactive mindset toward incident prevention.
  • Deep understanding of Linux systems, networking, and common troubleshooting practices.
  • Solid understanding of the network stack (e.g., TCP/IP, VPN, etc.), cloud architectures (VPC, subnets, firewalls, load balancers), service mesh (e.g., Istio) and storage (e.g., S3, EBS, etc.).
  • Knowledge of zero-downtime deployment strategies, blue/green and canary releases.
  • Exposure to compliance standards such as SOC 2, ISO 27001, or HIPAA. FedRAMP experience is a big plus.
  • Experience with chaos engineering or resilience testing practices.
  • Excellent problem-solving skills, collaborative mindset, and a strong grasp of agile, iterative development.
  • Self-driven, highly organised, and capable of independently managing priorities.
  • Curiosity to learn new things and discover new technologies.
  • Strong communication, presentation, and team collaboration skills.
  • Excellent written and verbal skills in English.
  • Prior experience with any of the above-mentioned tools is a bonus, but not a must!
Additional Information

We are the pioneers and trailblazers of a global IT Market Category (DEX) that is shaping the future of how the world works, giving our customers IT Teams total digital visibility across their enterprise. Our innovative solutions integrate real-time analytics, automation, and employee feedback across all endpoints. This enables our IT teams to solve complex technical challenges, create ever more productive workplaces, and deliver happy, satisfied employees in the digital workplace.

With over 1000 employees across 5 continents, Nexthink operates as One Team, connecting, collaborating and innovating to continuously grow. We call our employees 'Nexthinkers', and our commitment to diversity, inclusion, and equity is second to none. We currently have over 75 nationalities working with us, from all cultures and backgrounds, speaking many different languages.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Backend Software Engineer (Java)
Senior Backend Software Engineer (Java)

Nexthink • Bengaluru

Hybrid
INR 3,000,000 - 6,000,000
Health insurance
Hybrid work model
Flexible hours
+5
Professional Service Consultant
Professional Service Consultant

Nexthink India Digital Experience • Bengaluru

Hybrid
INR 1,800,000 - 2,800,000
Total rewards package
Software Backend Engineer
Software Backend Engineer

Nexthink • Karnataka

Hybrid
INR 2,600,000 - 4,200,000
Health insurance
Hybrid work model
Flexible hours & unlimited vacation
+4
Professional Services Consultant
Professional Services Consultant

Nexthink • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Health insurance
Hybrid work model
Unlimited vacation
+2
Site Reliability Engineer
Site Reliability Engineer

TVS Next • Chennai District

Hybrid
INR 1,200,000 - 2,000,000
Hybrid work model
Family health insurance coverage
Accelerated career paths
Platform Software Engineer - Core Infrastructure
Platform Software Engineer - Core Infrastructure

Nexthink • Bengaluru

Hybrid
INR 3,500,000 - 5,500,000
Health insurance through Bajaj
Hybrid work model
Unlimited vacation
+2
Tech lead
Tech lead

Nexthink • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Health insurance
Hybrid work model
Flexible hours
+3
Site Reliability Engineer
Site Reliability Engineer

Nice • Pune District

Hybrid
INR 1,800,000 - 2,800,000
NiCE-FLEX hybrid model
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

NICE • Pune District

Hybrid
INR 1,400,000 - 2,200,000
NICE-FLEX hybrid model
Senior Backend Software Engineer
Senior Backend Software Engineer

Nexthink • Bengaluru

Hybrid
INR 1,800,000 - 3,000,000
Health insurance
Hybrid work model
Unlimited vacation
+1