Senior Cloud Site Reliability Engineer

NICE

United States

Hybrid

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

NICE-FLEX hybrid

Job summary

NICE is seeking a seasoned Platform Reliability Engineer to monitor and manage production environments, build automation, and improve reliability across a suite of distributed applications. You will partner with development teams to enhance services, implement scalable design decisions, and drive efficient capacity planning.

Responsibilities include incident management, performance tuning, and creating sustainable systems through automation.

Qualifications

  • 3–6 years of experience in a similar role focusing on systems engineering, automation and reliability.
  • Proficiency in at least one programming language (e.g., Python, Go, Java, C#) and scripting (e.g., Bash, PowerShell).
  • Deep understanding of cloud platforms (AWS) and services (EC2, ECS, Lambda, DynamoDB).
  • Experience with infrastructure as code tools (CloudFormation, Terraform).
  • Strong CI/CD knowledge and experience with Jenkins, GitLab CI/CD or CircleCI.
  • Solid containerization knowledge (Docker, Kubernetes) and microservices.
  • Experience with monitoring/observability tools (Prometheus, Grafana, ELK, CloudWatch).
  • Experience with incident management and postmortems in cross-functional teams.

Responsibilities

  • Run production environment by monitoring availability and system health.
  • Build software and systems to manage platform infrastructure and applications.
  • Improve reliability, quality, and time-to-market of software solutions.
  • Measure and optimize system performance and drive innovation to meet customer needs.
  • Provide primary operational support and engineering for distributed applications.

Skills

Python
Go
Java
C#
Bash
PowerShell
AWS
Terraform
CloudFormation
CI/CD
Jenkins
GitLab CI/CD
CircleCI
Docker
Kubernetes
Prometheus
Grafana
ELK stack
CloudWatch

Tools

CloudFormation
Terraform
Jenkins
GitLab CI/CD
CircleCI
Ansible
Puppet
Chef

Job description

So, what's the role all about?
  • Run the production environment by monitoring availability and taking a holistic view of system health
  • Build software and systems to manage platform infrastructure and applications
  • Improve reliability, quality, and time-to-market of our suite of software solutions
  • Measure and optimize system performance, with an eye toward pushing our capabilities forward, getting ahead of customer needs, and innovating to continually improve
  • Provide primary operational support and engineering for multiple large distributed software applications
How will you make an impact?
  • Gather and analyze metrics from both operating systems and applications to assist in performance tuning and fault finding
  • Partner with development teams to improve services through rigorous testing and release procedures
  • Participate in system design consulting, platform management, and capacity planning
  • Create sustainable systems and services through automation and uplifts
  • Balance feature development speed and reliability with well-defined service level objectives
Have you got what it takes?
  • 3-6 years of working experience in a similar role, with a focus on systems engineering, automation, and reliability.
  • Proficiency in at least one programming language (e.g., Python, Go, Java, C#) and experience with scripting languages (e.g., Bash, PowerShell).
  • Deep understanding of cloud computing platforms (e.g., AWS), the working and reliability constraints of some of the prominent services (e.g., EC2, ECS, Lambda, DynamoDB etc)
  • Experience with infrastructure as code tools such as CloudFormation, Terraform.
  • Deep understanding of CI/CD concepts and experience with CI/CD tools such as Jenkins, GitLab CI/CD, or CircleCI.
  • Strong knowledge of containerization technologies (e.g., Docker, Kubernetes) and microservices architecture.
  • Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack, Cloudwatch).
  • Excellent problem-solving skills and the ability to troubleshoot complex issues in distributed systems.
  • Experience of Incident management and blameless postmortems that includes driving the incident response efforts during outages and other critical incidents, resolution, and communication in a cross-functional team setup.
You will have an advantage if you also have:
  • Handson experience of working with large Kubernetes Cluster. Certification will be an added plus.
  • Working experience of Grafana Observability Suite (Loki, Mimir, Tempo).
  • Administration and/or development experience of standard monitoring and automation tools such as Splunk, Datadog, Pagerduty Rundeck.
  • Familiarity with configuration management tools like Ansible, Puppet, or Chef.
  • Certifications such as AWS Certified DevOps Engineer, Google Cloud Professional DevOps Engineer, or equivalent.
Personal attributes:
  • Strong communication skills and the ability to collaborate effectively with cross-functional teams.
  • Team player - ability to work well in a close team environment.
  • Fast learner with ability to educate her/himself on relevant technologies
  • Ability to multitask and prioritize work
  • Ability to remain focused and calm under pressure
Enjoy NICE-FLEX!

At NICE, we work according to the NICE-FLEX hybrid model, which enables maximum flexibility: 2 days working from the office and 3 days of remote work, each week. Naturally, office days focus on face-to-face meetings, where teamwork and collaborative thinking generate innovation, new ideas, and a vibrant, interactive atmosphere.

Requisition ID:

Reporting into: Director, Network Operations.

Role Type: Individual Contributor.

#LI-Hybrid

About NiCE

NICELtd. (NASDAQ: NICE)software products are used by 25,000+ global businesses, including 85 of the Fortune 100 corporations, to deliver extraordinary customer experiences,fight financial crimeand ensure public safety.Every day, NiCE software managesmore than120 million customer interactions and monitors3+billion financial transactions.

Known as an innovation powerhouse that excels in AI, cloud and digital, NiCE is consistently recognized as the market leader in its domains, with over 8,500 employees across 30+ countries.

NiCE is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, age, sex, marital status, ancestry, neurotype, physical or mental disability, veteran status, gender identity, sexual orientation or any other category protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

NiCE • United States

Hybrid
GBP 70,000 - 100,000
NICE-FLEX hybrid model
Tech Manager, CX
Tech Manager, CX

NICE • United States

Hybrid
USD 150,000 - 230,000
NICE-FLEX hybrid model
DevOps Engineer
DevOps Engineer

NICE • United States

Hybrid
USD 120,000 - 180,000
NICE-FLEX hybrid work model
Global team exposure
Cloud Systems Administrator
Cloud Systems Administrator

Cisco Systems • Seattle (WA)

Hybrid
USD 110,000 - 160,000
NICE-FLEX hybrid work model
Senior AI Software Engineer
Senior AI Software Engineer

NICE • Seattle (WA)

On-site
USD 180,000 - 240,000
Principal Cloud Architect
Principal Cloud Architect

NiCE • Sandy (UT)

On-site
USD 180,000 - 260,000
Lead DevOps Engineer- Late Shift(2 PM - 10PM)
Lead DevOps Engineer- Late Shift(2 PM - 10PM)

NICE • United States

Hybrid
USD 100,000 - 130,000
Flexible working environment
Career growth opportunities
Principal Cloud Architect
Principal Cloud Architect

Socket.dev • Sandy (UT)

Hybrid
USD 180,000 - 240,000
Competitive pay
Benefits package
Professional development
+1
Senior Full Stack Software Engineer
Senior Full Stack Software Engineer

NICE • Sandy (UT)

On-site
USD 150,000 - 190,000
Technical Support Engineer
Technical Support Engineer

NICE • United States

Hybrid
USD 70,000 - 100,000