Staff SRE: Cloud-Native Reliability & SecOps Leader

Crunchyroll, LLC

Los Angeles (CA)

On-site

USD 211,000 - 263,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Salary plus performance bonus
Flexible time off
Medical, dental, vision insurance
401(k) plan with employer match
Commuter benefit

Job summary

Crunchyroll, LLC is seeking a Staff Site Reliability Engineer to join the Center for Data & Insights in the US. You will help design and operate resilient cloud-native data platforms powering Crunchyroll's consumer experiences, collaborating with Engineering, Data, Infrastructure and Security teams.

You will drive modern SRE practices (SLIs/SLOs, error budgets), automate operations, and strengthen SecOps by improving cloud/Kubernetes security, disaster recovery, and incident management across

Qualifications

  • 12+ years of experience in Site Reliability Engineering or related fields.
  • Deep expertise in Kubernetes and GCP for large-scale cloud-native platforms.
  • Strong IaC experience, preferably Terraform, with a focus on automation.
  • Solid Linux systems administration, networking, and distributed systems knowledge.
  • Proficiency in programming languages such as Go, Python, Java, or Shell.
  • Hands-on experience with modern observability platforms (Prometheus, Grafana, OpenTelemetry, Datadog).
  • Experience with incident management, reliability metrics (SLIs/SLOs), and capacity planning.
  • Knowledge of cloud and platform security, container security, and secure operations.

Responsibilities

  • Define and measure reliability, availability, and performance using SLIs, SLOs, and error budgets.
  • Establish and drive best practices for incident management and postmortems.
  • Build and evolve monitoring, logging, tracing, and alerting systems for proactive issue detection.
  • Identify operational inefficiencies and develop automation and self-service capabilities.
  • Design and optimize cloud-native infrastructure for scalable performance and cost efficiency.
  • Drive IaC, standardization, and deployment automation to improve reliability.
  • Lead capacity planning and performance optimization to scale platforms.
  • Develop disaster recovery, backup, and business continuity strategies.

Skills

Kubernetes
GCP
Terraform
Linux
Go
Python
Java
Shell
Observability
Prometheus
Grafana
OpenTelemetry
Datadog
SLIs/SLOs
Incident Mgmt

Tools

Docker

Job description

Crunchyroll, LLC is seeking a Staff Site Reliability Engineer to join the Center for Data & Insights in the US. You will help design and operate resilient cloud-native data platforms powering Crunchyroll's consumer experiences, collaborating with Engineering, Data, Infrastructure and Security teams.

You will drive modern SRE practices (SLIs/SLOs, error budgets), automate operations, and strengthen SecOps by improving cloud/Kubernetes security, disaster recovery, and incident management across

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: Cloud-Native Reliability & Security
Senior SRE: Cloud-Native Reliability & Security

Ellation, Inc. • Los Angeles (CA)

On-site
USD 211,000 - 263,000
Salary plus performance bonus
Flexible time off
Medical, dental, vision insurance
+2
Senior Cloud-Native SRE & Platform Reliability Lead
Senior Cloud-Native SRE & Platform Reliability Lead

Engg • Los Angeles (CA)

On-site
USD 180,000 - 240,000
Staff SRE – Cloud-Native Reliability & SecOps
Staff SRE – Cloud-Native Reliability & SecOps

Crunchyroll • San Francisco (CA)

On-site
USD 233,000 - 292,000
Performance bonus
Flexible time off
Medical insurance
+5
Senior SRE — Cloud Data Platform & SecOps
Senior SRE — Cloud Data Platform & SecOps

Engg • San Francisco (CA)

On-site
USD 180,000 - 250,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Engg • Los Angeles (CA)

On-site
USD 180,000 - 240,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Engg • San Francisco (CA)

On-site
USD 180,000 - 250,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Ellation, Inc. • Los Angeles (CA)

On-site
USD 211,000 - 263,000
Salary plus performance bonus
Flexible time off
Medical, dental, vision insurance
+2
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Crunchyroll • San Francisco (CA)

On-site
USD 233,000 - 292,000
Performance bonus
Flexible time off
Medical insurance
+5
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Crunchyroll, LLC • Los Angeles (CA)

On-site
USD 211,000 - 263,000
Salary plus performance bonus
Flexible time off
Medical, dental, vision insurance
+2
Staff SRE: Scale, Observability & Automation Leader
Staff SRE: Scale, Observability & Automation Leader

Replit • Northern (KY)

Hybrid
USD 180,000 - 260,000
Salary & equity
401(k) matching
Health, dental, vision, life
+9