Specialist Cloud Site Reliability Engineer

NICE

Pune District

On-site

INR 450,000 - 650,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

NICE-FLEX hybrid work model

Job summary

NiCE is seeking an experienced Site Reliability Engineer to strengthen reliability, automation, and observability across hybrid cloud environments. You’ll own uptime, latency, SLOs, and incident response with emphasis on scalable infrastructure.

You will implement IaC with Terraform and Helm, improve CI/CD with Jenkins and GitHub Actions, and mentor junior engineers. Candidates should have 8+ years in Kubernetes/EKS/ECS and strong AWS and OpenTelemetry/Grafana skills.

Qualifications

  • 8+ years of production experience with Kubernetes, EKS and ECS in distributed environments.
  • Expertise in AWS services: EC2, Lambda, IAM, RDS, S3, ALB/NLB, VPC, PrivateLink.
  • Proficiency with Terraform, Helm, Jenkins, GitHub Actions, and GitOps (ArgoCD or Flux).
  • Strong observability knowledge: metrics, logs, traces, and distributed monitoring.
  • Hands-on with Prometheus, Grafana, Loki, Tempo, OpenTelemetry, or equivalent tools.
  • Strong Linux, networking fundamentals, and system performance tuning.
  • Proficiency in Python, Go, or Shell scripting for automation.
  • Experience in incident response, RCA, and on-call operations.

Responsibilities

  • Design and implement scalable, reliable, and resilient systems across hybrid or multi-cloud environments (AWS/EKS/ECS).
  • Drive improvements in system uptime, latency, and service health metrics (SLOs, SLIs, SLAs).
  • Build and manage infrastructure automation using Terraform, Helm, and Kubernetes; enhance CI/CD pipelines with Jenkins and GitHub Actions.
  • Own and enhance the observability stack (Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Mimir).
  • Define and implement SLOs and error budgets; enable teams to monitor reliability metrics.
  • Lead major incident response, root cause analysis (RCA), and blameless postmortems; ensure operational readiness before releases.
  • Ensure platform security and compliance; collaborate with InfoSec and compliance teams.
  • Mentor junior SREs and developers on reliability practices, automation, and observability.

Skills

Kubernetes
AWS
GitOps
Terraform
Helm
Jenkins
GitHub Actions
OpenTelemetry
Monitoring
On-call

Tools

ArgoCD
Flux
Prometheus
Grafana
Loki
Tempo
OpenTelemetry
Mimir

Job description

So, what's the role all about?

At NiCE, we don't limit our challenges. We challenge our limits. Always. We're ambitious. We're game changers. And we play to win. We set the highest standards and execute beyond them. And if you're like us, we can offer you the ultimate career opportunity that will light a fire within you.

How will you make an impact?
Reliability & Performance
  • Design and implement scalable, reliable, and resilient systems across hybrid or multi-cloud environments (primarily AWS/EKS/ECS/Lambda)
  • Drive improvements in system uptime, latency, and overall service health metrics (SLOs, SLIs, SLAs)
Automation & Infrastructure as Code
  • Build and manage infrastructure automation using Terraform, Helm, and Kubernetes
  • Improve CI/CD pipelines using Jenkins, GitHub Actions, ensuring safe and automated rollouts, monitoring, and rollbacks
Observability & Monitoring
  • Own and enhance the observability stack (Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Mimir, etc.)
  • Define and implement SLOs and error budgets; enable development teams to monitor and act on reliability metrics
Incident Management
  • Lead major incident response, root cause analysis (RCA), and blameless postmortems
  • Partner with product teams to define and enforce operational readiness standards before production releases
Security & Compliance
  • Ensure platform-level security and compliance with organizational and regulatory standards
  • Collaborate with InfoSec and compliance teams to maintain a secure and auditable infrastructure
Technical Leadership
  • Mentor junior SREs and developers on reliability practices, automation, and observability
  • Contribute to technical roadmaps and reliability-focused design reviews
Have you got what it takes?
  • 8+ Years Strong experience with Kubernetes, EKS, ECS and containerized workloads in production
  • Expertise in AWS services (EC2, Lambda, IAM, RDS, S3, ALB/NLB, VPC, PrivateLink, etc.)
  • Proficiency with Terraform, Helm, Jenkins, GitHub Actions, and GitOps (ArgoCD or Flux)
  • Deep understanding of observability frameworks - metrics, logs, traces, and distributed monitoring
  • Hands-on experience with Prometheus, Grafana, Loki, Tempo, Alloy, OpenTelemetry, or equivalent tools
  • Strong knowledge of Linux, networking fundamentals, and system performance tuning
  • Familiarity with Python, Go, or Shell scripting for automation and custom tooling
  • Practical experience in incident response, RCA, and on-call operations
What's in it for you?

Join an ever-growing, market disrupting, global company where the teams - comprised of the best of the best - work in a fast-paced, collaborative, and creative environment! As the market leader, every day at NICE is a chance to learn and grow, and there are endless internal career opportunities across multiple roles, disciplines, domains, and locations.

Enjoy NICE-FLEX!

At NICE, we work according to the NICE-FLEX hybrid model, which enables maximum flexibility: 2 days working from the office and 3 days of remote work, each week.

Requisition ID - 11465

Reporting into: Tech Manager

Role Type: Individual Contributor

About NiCE

NICELtd. (NASDAQ: NICE)software products are used by 25,000+ global businesses, including 85 of the Fortune 100 corporations, to deliver extraordinary customer experiences,fight financial crimeand ensure public safety.Every day, NiCE software managesmore than120 million customer interactions and monitors3+billion financial transactions.

Known as an innovation powerhouse that excels in AI, cloud and digital, NiCE is consistently recognized as the market leader in its domains, with over 8,500 employees across 30+ countries.

NiCE is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, age, sex, marital status, ancestry, neurotype, physical or mental disability, veteran status, gender identity, sexual orientation or any other category protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

NICE • Pune District

Hybrid
INR 1,400,000 - 2,200,000
NICE-FLEX hybrid model
Senior Technical Support Engineer
Senior Technical Support Engineer

Nice • Pune District

On-site
INR 900,000 - 1,500,000
NICE-FLEX
Hybrid work model
Senior Technical Support Engineer
Senior Technical Support Engineer

Nice Ltd. • Pune District

Hybrid
INR 1,200,000 - 2,400,000
Hybrid work model
Specialist Software Engineer, CX
Specialist Software Engineer, CX

Satmetrix Systems, Inc. • Pune District

Hybrid
INR 3,000,000 - 4,200,000
NiCE-FLEX hybrid model
Site Reliability Engineer
Site Reliability Engineer

Nice • Pune District

Hybrid
INR 1,800,000 - 2,800,000
NiCE-FLEX hybrid model
Specialist DevOps Engineer
Specialist DevOps Engineer

Nice • Pune District

Hybrid
INR 1,200,000 - 2,400,000
Senior Specialist DevOps Engineer
Senior Specialist DevOps Engineer

Nice Ltd. • Pune District

Hybrid
INR 2,400,000 - 4,000,000
Hybrid work model
Global exposure
Career growth opportunities
Tech Manager, Engineering ( .NET , AWS , AI)
Tech Manager, Engineering ( .NET , AWS , AI)

NICE • Pune District

Hybrid
INR 4,200,000 - 6,800,000
Senior Software Engineer (Dot Net, AWS, AI)
Senior Software Engineer (Dot Net, AWS, AI)

Livevox Solutions • Pune District

Hybrid
INR 1,500,000 - 2,300,000
NICE-FLEX Hybrid Model
Tech Manager, Cloud Operations, Actimize
Tech Manager, Cloud Operations, Actimize

Nice • Pune District

Hybrid
INR 4,500,000 - 7,500,000
NiCE-FLEX hybrid model