Senior Site Reliability Engineer

Duetto

Las Vegas (NV)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Duetto is looking for a Senior Site Reliability Engineer in Las Vegas, NV to enhance our infrastructure for the hospitality industry. This role demands 5+ years in SRE, Ops, or DevOps, proficiency in AWS, and a collaborative spirit. Responsibilities include system architecture, tool development, and ensuring high uptime and security. If you're passionate about technology and have a track record in reliable systems, apply to join our innovative team.

Qualifications

  • 5+ years of experience in an Ops, DevOps, or SRE role.
  • Experience in System Design and Architecture.
  • Engineer-level understanding of networking and security concepts.
  • Experience with AWS ecosystem tools including IAM, VPC, EC2, RDS, etc.
  • Proven ability to troubleshoot and resolve complex incidents.

Responsibilities

  • Architect and implement AWS infrastructure solutions.
  • Design and maintain tools for SaaS product operations.
  • Partner with developers to enhance system reliability and performance.
  • Lead efforts to ensure systems are secure by default.
  • Participate in weekly on-call rotation.

Skills

AWS
Java
Python
DevOps
Infrastructure as Code (Terraform)
Security compliance
Prometheus
CI/CD Tools
Networking

Tools

GitHub
Jenkins
Chef
DataDog
ECS/EKS

Job description

Duetto, the industry-leading hospitality revenue management system, leads the way in helping hotels, resorts and casinos optimize revenue and boost profit. Our leading SaaS platform, expanding suite of products, and incredibly skilled team have been at the heart of our continued success and our ambition for future growth knows no bounds.

Duetto is building the future of hotel revenue strategy. We’re not just another SaaS company — we’re redefining what’s possible for hotels through our category-creating platform, the Revenue & Profit Operating System.

Role Summary / Purpose

We are seeking a highly experienced Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have a proven track record of designing, implementing, and maintaining scalable, secure, and highly reliable systems. As a key contributor, you will collaborate with cross-functional teams to drive architecture decisions, implement best practices, and ensure high system availability.

Our technology stack is built on AWS and primarily consists of:

  • Java
  • Python
  • NoSql
  • Single-page JavaScript web techniques (jQuery, Backbone, React, and RequireJS)
  • Patent-pending analytical methods on top of MongoDB
  • Postgres
  • Terraform/Terragrunt and Chef for IaC
  • DataDog and Prometheus
  • GitHub for source control
  • GitHub Actions and Jenkins for CI/CD
Key Responsibilities
  • Architect and implement infrastructure solutions to facilitate seamless migration of critical systems while ensuring uptime, reliability, and a high-quality experience for end users.
  • Design, develop, test, and maintain tools and processes to efficiently manage and operate SaaS products hosted on AWS, with a focus on scalability and automation.
  • Partner with developers to enhance the reliability, performance, scalability, and security of server and application architectures.
  • Build and maintain critical components of our infrastructure, emphasizing robustness, security, and high availability to meet demanding service-level expectations.
  • Foster strong cross-team collaboration by driving engagement, promoting shared goals, and ensuring alignment across technical and non-technical teams.
  • Lead efforts to ensure systems are secure by default, addressing vulnerabilities proactively and implementing best practices for cybersecurity preparedness.
  • Be willing to learn and adopt AI in DevOps/SRE workflows.
  • Be the last line of support for services that thousands of customers (hotels, resorts, casinos, etc.) around the world depend on 24/7.
  • Troubleshoot on-call incidents to ensure rapid resolution and minimal service disruption. Participate in detailed Root Cause Analysis (RCA) to identify underlying issues and work cross-functionally to implement preventative measures and long-term solutions, ensuring similar problems are avoided in the future.
Qualifications
Required Qualifications
  • 5+ years of experience in an Ops, DevOps or SRE role.
  • Experience in System Design and Architecture.
  • Engineer-level experience with networking and security concepts.
  • Understanding of fundamentals behind load balancing technologies. Experience configuring Layer 7 load-balancing is a plus.
  • Experience collaborating with engineers on architecture decisions.
  • Experience administering Cloud Computing Services such as AWS (preferred), Azure, or GCP, including working knowledge of permissions structures, multi-account management structures, and single sign-on(SO).
  • Experience with AWS ecosystem tools such as AWS IAM, VPC, EC2, ELB, RDS, S3, Lambda, API Gateway, Secrets Manager, KMS, CloudWatch, CloudTrail.
  • Experience with security compliance certifications such as SOC2.
  • Experience working in an environment with a heavy emphasis on DevOps and Service Reliability mindset.
  • Experience provisioning, configuring, administering, and using enterprise monitoring ecosystems like Prometheus, Grafana, DataDog or similar.
  • Experience with CI/CD Tools such as GitHub, GitHub Actions, JFrog Artifactory, Jenkins, and GitOps methodologies.
  • Experience using and writing infrastructure-as-code using Terraform.
  • Experience with configuration-management toolsets such as Chef or Puppet.
  • Experience with containers and container orchestration tools such as ECS/EKS (a plus).
  • Experience managing infrastructure and contributing as part of a multi-user infrastructure team, using Terraform and associated toolsets. Relevant SOC2 experience is also a plus.
  • Fluency in reading Java, Ruby, Bash/Zsh, HCL, Python and Javascript.
  • Strong experience in troubleshooting and resolving complex on-call incidents with a focus on minimizing service disruption and downtime.
  • Proven ability to lead and participate in detailed Root Cause Analysis (RCA) processes to identify and address underlying issues effectively.
  • Demonstrated expertise in implementing preventative measures and long-term solutions based on RCA findings to ensure recurring issues are mitigated.
  • Experience constructing and maintaining build/deploy automation tooling.
  • Participate in weekly on-call rotation.
  • Ability to work both independently and within a team environment.
  • A passion for technology with a drive to stay up to date with technology and best practices.
Ideal Candidate
  • Team Player - Works well with others, highly collaborative and acts as a strong partner to other team members and functions.
  • Execution - Desire to work on a fast paced team and help set direction and architecture.
  • Creativity - Thrives in an environment without a set playbook.
  • Quality - Takes pride in delivering robust and high quality implementations.
  • Ownership - Enjoys owning and driving projects.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Duetto • United States

On-site
USD 120,000 - 160,000
Senior SRE: Scale, Secure, and Automate SaaS
Senior SRE: Scale, Secure, and Automate SaaS

Duetto • United States

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

R&D • United States

On-site
USD 90,000 - 140,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Senior SRE: Scale, Secure, and Automate SaaS
Senior SRE: Scale, Secure, and Automate SaaS

Duetto • Las Vegas (NV)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Cvent, Inc. • Tysons (VA)

Hybrid
USD 100,000 - 130,000