Senior Site Reliability Engineer

GoGuardian

El Segundo (CA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

GoGuardian is seeking a Senior Site Reliability Engineer to design, scale, and maintain cloud infrastructure powering our core products. You will collaborate with engineering teams to drive reliability, observability, and secure dev tooling across production environments.

The role sits in Tech Foundation, focusing on shared data services, cloud infra, and tooling to empower product teams to deliver software efficiently and securely.

Qualifications

  • 5+ years of SRE/DevOps experience supporting production SaaS.
  • Strong AWS, Kubernetes, Terraform and CI/CD experience.
  • Experience with data layers such as MongoDB, Redshift, OpenSearch.
  • Linux fundamentals and shell scripting proficiency.
  • Ability to read and debug JS/TS, Python or Go code.

Responsibilities

  • Architect and maintain scalable, secure cloud infrastructure for high availability.
  • Enhance observability and monitoring to reduce noise and improve detection.
  • Lead on-call rotations and incident post-mortems to drive improvements.
  • Modernise deployment pipelines and automation to speed delivery.
  • Collaborate with product teams to promote reliability best practices.
  • Ensure security standards and compliance across cloud infra.

Skills

AWS core services
Terraform (IaC)
Kubernetes (EKS)
CI/CD pipelines
Linux scripting
Programming languages (JS/TS, Python,/
Observability/Monitoring
On-call / incident response
Communication & collaboration

Tools

Jenkins
GitHub Actions
AWS CodeBuild/CodePipeline
Terraform

Job description

We're looking for a Senior Site Reliability Engineer (SRE) to help design, scale, and maintain the infrastructure that powers our core products and services. In this role, you'll collaborate with engineering teams to drive operational excellence, optimise system performance, and ensure high availability across production environments. This position sits on Tech Foundation, a team that manages core cloud infrastructure, shared data services, and developer tooling to empower our product teams to deliver software efficiently and securely. The ideal candidate brings a strong background in cloud infrastructure, automation, and modern reliability practices, with a passion for solving complex operational challenges in a collaborative environment.

Responsibilities:

  • Architect and maintain scalable, secure cloud infrastructure to ensure high availability for core products.
  • Enhance observability and monitoring frameworks to deliver highly accurate alerts, minimising noise and improving incident detection.
  • Participate in on-call rotations and lead incident response, ensuring comprehensive post-mortems and RCAs are completed to drive systemic improvements.
  • Optimise and modernise deployment pipelines and automation workflows to maximise engineering velocity and operational safety.
  • Partner with product development teams to provide infrastructure support, review architectural changes, and promote reliability best practices.
  • Implement and uphold robust security standards and compliance controls across all managed cloud infrastructure.

Requirements:

  • 5+ years of professional experience in Site Reliability Engineering, Infrastructure, or DevOps roles supporting production SaaS applications.
  • Strong proficiency with AWS core services (including EC2 VPC, S3) along with experience in Serverless frameworks and managed Kubernetes environments like EKS.
  • Extensive experience writing and managing Infrastructure as Code (IaC) using Terraform.
  • Familiarity with configuring, troubleshooting, and maintaining data layers such as MongoDB, Redshift, and OpenSearch.
  • Experience with GCP environments or technologies like Firestore is a plus.
  • Experience managing or modernising CI/CD pipelines and deployment workflows utilising systems like Jenkins, AWS CodeBuild/CodePipeline, or GitHub Actions.
  • Deep understanding of Linux operating system fundamentals and Unix shell scripting.
  • Ability to read and debug code written in JavaScript/TypeScript, Python, or Go to effectively troubleshoot underlying service errors.
  • Strong communication and collaboration skills, with a track record of driving technical decisions and establishing team-wide operational standards.
  • Eager to take initiative in a fast-paced, ever-changing, dynamic environment.
  • Fueled by the opportunity to truly impact the education landscape.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Phoenix (AZ)

On-site
USD 120,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Engg • Tempe (AZ)

On-site
USD 140,000 - 190,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

On-site
USD 150,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000