Senior Site Reliability Expert

Upserve

United States

On-site

USD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Upserve is seeking an experienced Site Reliability Engineer to design, operate, and improve the reliability of our restaurant-technology platform. You will automate provisioning, monitoring, and incident response in collaboration with software engineers, QA, and product teams.

The role emphasizes cloud-based architecture, cost-optimized infrastructure, and adherence to IaC and DevOps practices. Remote-friendly environment with opportunity for growth across technical and leadership paths.

Qualifications

  • Strong knowledge of AWS and cloud infrastructure best practices.
  • Hands-on experience with Docker, Kubernetes, and Linux systems.
  • Experience with configuration management tools (Chef, Puppet, Ansible, Salt).
  • Infrastructure as code: Terraform/OpenTofu, automation and scripting.

Responsibilities

  • Design, operate and ensure reliability of Upserve product infrastructure.
  • Automate and monitor systems to support product development.
  • Architect scalable, cost-efficient cloud infrastructure with high availability.
  • Collaborate with developers, QA, PMs and other teams.
  • Adhere to best practices: IaC, monitoring, security, SRE/DevOps methods.
  • Provide timely remediation during production incidents and be on-call as needed.

Skills

AWS
Docker
Kubernetes
Linux
Chef-Puppet-Ansible-Salt
Terraform-OpenTofu
Shell scripting
Python-Ruby-Go
Agile-CD

Tools

Terraform
OpenTofu

Job description

About Upserve, Inc.

Upserve is an established restaurant technology and payments platform entering a restart phase under private equity ownership. The company serves independent and multi-location full-service restaurants and operates across software, payments, and financial services.

Role Summary

Our SRE team is responsible for the design, operation and reliability of Upserve’s product infrastructure. We collaborate with teams across the company to make this happen: Developers, QA, PMs, etc.

Key Responsibilities
  • Initiate and contribute to continuous improvement of our software delivery processes and practices in a multi-location, multidisciplinary team to empower and accelerate product development
  • Use automation extensively to design, configure, manage, and monitor systems in support of our product development teams
  • Design and architect operational solutions with the specific goal of increasing the standardization, automation, repeatability, cost-efficiency and consistency of operational tasks
  • Working with developers and other SREs to design and build scalable, reliable and cost-efficient Cloud infrastructure
  • Adhere to and advocate for best practices, including Infrastructure as Code, monitoring, high availability, disaster recovery, security, and SRE/DevOps methodologies
  • Provide timely assistance and remediation solutions during critical situations and production incidents to help resolve service problems (You will be on call for periods of time)
Required Qualifications
  • Strong knowledge of Amazon Web Services
  • Strong experience with Docker, Kubernetes & Linux Systems
  • Experience with configuration management tools such as Chef, Puppet, Ansible, Salt
  • Experience with Infrastructure as code practices: we use Terraform & OpenTofu
  • Ability to read & write complex scripts using Shell
  • Ability to read & understand programming languages: Python, Ruby, Go, etc.
  • Good understanding of Agile development and continuous delivery best practices, software engineering tools, processes, methods and testing
  • Ability to collaborate effectively with other teams
  • Ability to plan, organize, prioritize and stay focused
  • Good experience provisioning and managing infrastructures with high availability constraints
  • Good experience with cloud cost optimization
First 90 Days: Success Outcomes
  • You are a problem solver who does not shy away from tackling complexity and critical thinking
  • You have a strong will to learn, grow and get out of your comfort zone
  • You have great energy and passion for technology
  • You are able to express yourself flawlessly in English
  • You have strong interpersonal skills
Opportunity
  • Lots of autonomy, flexible work culture and possibility of remote work
  • Development of high traffic products, used at the global scale
  • Exposure to modern and proven technology
  • Opportunity to learn and expand your skill set
  • Tons of growth opportunities into technical or people management roles
  • Opportunity to join a fast-paced, high-growth company
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE – Remote, High-Impact Cloud Infra
Senior SRE – Remote, High-Impact Cloud Infra

Upserve • United States

On-site
USD 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Senior Software Engineer, Site Reliability
Senior Software Engineer, Site Reliability

Upstart • United States

Hybrid
USD 166,000 - 231,000
Competitive compensation
401(k) matching
Employee Stock Purchase Plan
+5
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Supio • San Francisco (CA)

On-site
USD 170,000 - 220,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
SRE Leader
SRE Leader

Kontakt Micro-Location Sp. Z.o.o. • New York (NY)

Hybrid
USD 180,000 - 260,000
Equity in a high-growth company
Health, dental, and vision coverage
401k
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Duetto • Las Vegas (NV)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Duetto • United States

On-site
USD 120,000 - 160,000