Staff Site Reliability Engineer

Engg

Los Angeles (CA)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Crunchyroll is seeking a Staff Site Reliability Engineer to join the Center for Data & Insights (CDI) in the US. You will drive reliability, security, and scalability of data platforms, partnering with Engineering, Data, Infrastructure, Product, and Security teams.

The role emphasizes ownership, proactive automation, and modern SRE practices such as SLIs, SLOs and error budgets, with mentorship and cross-functional leadership across CDI.

Qualifications

  • 12+ years in Site Reliability Engineering, Platform or Infra roles.
  • Kubernetes and GCP expertise at scale.
  • IaC with Terraform and strong automation focus.

Responsibilities

  • Define reliability strategies using SLIs, SLOs, and error budgets.
  • Lead observability, monitoring, and incident management practices.
  • Drive capacity planning, performance optimization, and disaster recovery.
  • Improve security controls and SecOps alignment with CDI platforms.
  • Mentor engineers and promote operational excellence across teams.

Skills

Kubernetes
GCP
Terraform
Linux
Go
Python
Java
Shell
Prometheus
Grafana
OpenTelemetry
Datadog
Incident Mgmt
SLIs/SLOs

Job description

About Crunchyroll

Founded by fans, Crunchyroll delivers the art and culture of anime to a passionate community. We super-serve over 100 million anime and manga fans across 200+ countries and territories, and help them connect with the stories and characters they crave. Whether that experience is online or in-person, streaming video, theatrical, games, merchandise, events and more, it’s powered by the anime content we all love. Join our team, and help us shape the future of anime!

About the role

We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the US and play a critical role in advancing the reliability, scalability, performance, and security of Crunchyroll's consumer-facing data platforms. As a senior technical leader, you will partner closely with Engineering, Data, Infrastructure, Product, and Security teams to design and operate resilient cloud-native systems that power critical business and customer experiences. You will drive initiatives across observability, incident management, automation, capacity planning, disaster recovery, and operational excellence while helping teams adopt modern SRE practices such as SLIs, SLOs, and error budgets. The ideal candidate combines deep expertise in large-scale distributed systems with a strong sense of ownership, collaboration, and service leadership. You are passionate about building highly reliable platforms, eliminating operational toil through automation, and enabling engineering teams to move quickly and safely. In addition, you will champion SecOps best practices by driving vulnerability management, supporting penetration testing initiatives, improving security observability, strengthening cloud and Kubernetes security controls, and ensuring operational readiness for emerging threats. This is a unique opportunity to shape reliability and security engineering practices across CDI while helping build a world-class data and insights ecosystem that enables informed decision‑making throughout Crunchyroll.

Core Areas of Responsibility
  • Reliability Engineering : Define, measure, and continuously improve the reliability, availability, and performance of CDI platforms through SLIs, SLOs, and error budgets.
  • Operational Excellence : Establish and drive best practices for incident management, root cause analysis, postmortems, and service ownership across engineering teams.
  • Observability & Monitoring : Build and evolve comprehensive monitoring, logging, tracing, and alerting capabilities to enable proactive issue detection and rapid resolution.
  • Automation : Identify operational inefficiencies and develop automation, self-service capabilities, and self-healing mechanisms to improve engineering productivity.
  • Platform Scalability : Design and optimize cloud-native infrastructure and services to support growing business demands while maintaining performance and cost efficiency.
  • Infrastructure Engineering : Drive Infrastructure as Code (IaC), platform standardization, and deployment automation to improve consistency, reliability, and operational agility.
  • Capacity Planning & Performance : Lead capacity planning and performance optimization initiatives to ensure platforms can scale predictably and efficiently.
  • Disaster Recovery & Resilience : Develop and regularly validate disaster recovery, backup, and business continuity strategies to ensure platform resiliency.
  • Security Operations (SecOps) : Partner with Crunchyroll's security team to integrate security controls, operational risk management, and security best practices into platform operations and engineering workflows.
  • Vulnerability Management : Own the triage and remediation of identified vulnerabilities across infrastructure, platform, container, and application security vulnerabilities through established Crunchyroll vulnerability management processes.
  • Penetration Testing & Security Remediation : Support penetration test scoping activities by providing technical context on CDI platforms. Own the triage, prioritization, and remediation of resulting findings to drive timely resolution and strengthen platform security posture.
  • Cloud & Kubernetes Security : Implement and maintain secure cloud, container, and Kubernetes environments following least-privilege, defense-in-depth, and Zero Trust principles.
  • Cross-Functional Leadership : Collaborate with Engineering, Data, Product, Infrastructure, and Security teams to drive reliability, scalability, and security initiatives across CDI.
  • Mentorship & Engineering Excellence : Mentor engineers and champion a culture of operational excellence, reliability, ownership, continuous improvement, and security awareness.
About You
  • 12+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Engineering, or related disciplines, with a proven track record of operating and scaling production-critical systems.
  • Deep expertise in Kubernetes and GCP , including the design, deployment, and operation of highly available, cloud-native platforms at scale.
  • Strong Infrastructure as Code (IaC) experience , preferably with Terraform, and a commitment to automation, standardization, and operational efficiency.
  • Solid foundation in Linux systems administration, networking, and distributed systems , with the ability to troubleshoot complex production issues across multiple layers of the technology stack.
  • Proficiency in one or more programming and scripting languages , such as Go, Python, Java, or Shell, with a focus on automation and platform engineering.
  • Hands‑on experience with modern observability platforms and practices , including Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent monitoring and telemetry solutions.
  • Demonstrated expertise in incident management, service reliability, capacity planning, performance optimization, and operational excellence , including the implementation of SLIs, SLOs, and error budgets.
  • Strong understanding o
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Engg • San Francisco (CA)

On-site
USD 180,000 - 250,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Ellation, Inc. • Los Angeles (CA)

On-site
USD 211,000 - 263,000
Salary plus performance bonus
Flexible time off
Medical, dental, vision insurance
+2
Senior Cloud-Native SRE & Platform Reliability Lead
Senior Cloud-Native SRE & Platform Reliability Lead

Engg • Los Angeles (CA)

On-site
USD 180,000 - 240,000
Senior SRE — Cloud Data Platform & SecOps
Senior SRE — Cloud Data Platform & SecOps

Engg • San Francisco (CA)

On-site
USD 180,000 - 250,000
Staff Site Reliability Engineer, SRE
Staff Site Reliability Engineer, SRE

Jobtailor • California (MO)

On-site
USD 120,000 - 210,000
Senior SRE: Cloud-Native Reliability & Security
Senior SRE: Cloud-Native Reliability & Security

Ellation, Inc. • Los Angeles (CA)

On-site
USD 211,000 - 263,000
Salary plus performance bonus
Flexible time off
Medical, dental, vision insurance
+2
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000