Remote Senior Site Reliability Engineer — Scale & Observability

Epic for Kids

United States

On-site

USD 160,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Epic Kids is seeking a Senior Site Reliability Engineer to drive stability, observability, and reliability of our platform. You’ll own parts of our GCP infrastructure, container platform, CI/CD pipelines, and observability stack, partnering with product and data teams to keep applications running smoothly.

This role is fully remote, US-based, and involves on-call rotations and incident response. You’ll work with a global engineering team to raise reliability standards and reduce toil.

Qualifications

  • Bachelor's degree or higher in Computer Science, Software Engineering, or a related field.
  • 5+ years of experience in infrastructure, platform, DevOps, or a related engineering role, with a track record of measurably improving production reliability—including defining SLOs, reducing incident frequency or MTTR, and eliminating recurring failure modes.
  • Hands‑on experience with Google Cloud Platform (GCP), including GCE, GCS, VPC, IAM, Cloud Monitoring, and related services.

Responsibilities

  • Drive the reliability of Epic's infrastructure—set and track SLOs/SLIs, reduce toil, and engineer out recurring instability.
  • Build and operate the cloud infrastructure and container platform for high availability, scalability, and cost efficiency—including workload scheduling, autoscaling, networking, and graceful failure handling.
  • Maintain and improve CI/CD pipelines for fast, safe delivery across engineering teams.
  • Own and evolve the observability stack—metrics, logs, traces, dashboards, and alerts.
  • Manage infrastructure as code across the organization, with a focus on consistency, change safety, and reproducibility.
  • Own platform security practices—including secrets management, IAM policies, and network segmentation.
  • Support compliance‑aware infrastructure practices—including vulnerability management, access reviews, audit‑evidence flows, and incident‑response readiness.
  • Participate in a frequent on‑call rotation; drive incident response, blameless post‑mortems, and follow‑through on systemic fixes.
  • Partner with product and data engineering teams to troubleshoot platform issues and guide developers on infrastructure best practices.

Skills

SRE practices
Cloud infrastructure
Observability
CI/CD

Education

Bachelor's degree in Computer Science or related field

Tools

GCP
Docker
Kubernetes (GKE)
GitHub Actions
ArgoCD
Jenkins
Terraform
Python
Bash

Job description

Epic Kids is seeking a Senior Site Reliability Engineer to drive stability, observability, and reliability of our platform. You’ll own parts of our GCP infrastructure, container platform, CI/CD pipelines, and observability stack, partnering with product and data teams to keep applications running smoothly.

This role is fully remote, US-based, and involves on-call rotations and incident response. You’ll work with a global engineering team to raise reliability standards and reduce toil.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE — Remote Platform Reliability for Education
Senior SRE — Remote Platform Reliability for Education

TrulyHired • United States

On-site
USD 160,000 - 200,000
Senior SRE — Remote (US) | GCP, Kubernetes & CI/CD
Senior SRE — Remote (US) | GCP, Kubernetes & CI/CD

Epic Kids Inc. • United States

Remote
USD 160,000 - 200,000
Fully remote
Remote Senior SRE: GCP, Observability & Reliability Lead
Remote Senior SRE: GCP, Observability & Reliability Lead

United States Digital Space LLC • United States

Remote
USD 130,000 - 170,000
Remote Site Reliability Engineer - Cloud & Observability
Remote Site Reliability Engineer - Cloud & Observability

Prove • United States

On-site
USD 120,000 - 150,000
Competitive salary
Equity Plan
401(k) match
+2
Remote Site Reliability Engineer II - Observability
Remote Site Reliability Engineer II - Observability

QUEST DIAGNOSTICS INC • Secaucus (NJ)

Hybrid
USD 111,000 - 130,000
Medical, dental & vision insurance
401(k) with company match
Education assistance
+2
Senior Site Reliability Engineer – Remote, Impact & Automation
Senior Site Reliability Engineer – Remote, Impact & Automation

Midwest Startups • United States

On-site
USD 175,000 - 185,000
Market-leading medical, dental, and視on
Stock options
Premium-Tier Origin Financial Wellness
+6
Senior Site Reliability Engineer — Scale & Observability
Senior Site Reliability Engineer — Scale & Observability

Pivotal Health • New York (NY)

Hybrid
USD 230,000 - 260,000
Equity
Health, dental, vision
401(k)
+2
Senior Site Reliability Engineer – Scale & Observability
Senior Site Reliability Engineer – Scale & Observability

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Remote Senior DevOps Engineer: Observability & SRE
Remote Senior DevOps Engineer: Observability & SRE

United States Digital Space LLC • United States

Remote
USD 120,000 - 150,000
Health insurance
Flexible PTO
401K with employer match
+2
Remote Senior Site Reliability Engineer — Reliability Lead
Remote Senior Site Reliability Engineer — Reliability Lead

Priority Technology Holdings, Inc. • Alpharetta (GA)

On-site
USD 129,000 - 161,000
401(k) match
Employee Stock Purchase Program (ESPP)
Medical, dental, and vision coverage
+1