Staff Site Reliability Engineer

Crunchyroll, LLC

San Francisco (CA)

On-site

USD 233,000 - 292,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Salary plus performance bonus
Flexible time off
Medical insurance
Dental insurance
Vision insurance
STD insurance
LTD insurance
Life insurance
Health Savings Account
FSA - Health care and dependent care
401(k) with employer match
Parental/maternity/paternity support
Pet insurance
Pet-friendly offices

Job summary

Crunchyroll, LLC is seeking a Staff Site Reliability Engineer to advance the reliability, scalability, performance, and security of CDI's data platforms in the US. You will collaborate with Engineering, Data, Infrastructure, Product, and Security teams to design and operate resilient cloud-native systems and drive SRE practices across CDI.

The role emphasizes ownership, automation, capacity planning, disaster recovery, and SecOps integration, with opportunities to mentor engineers and shape

Qualifications

  • 12+ years of experience in SRE/Platform/Infra engineering with production systems.
  • Deep Kubernetes and GCP expertise at scale.
  • Strong IaC experience, preferably Terraform.
  • Solid Linux, networking, and distributed systems fundamentals.
  • Proficient in Go, Python, Java, or Shell for automation.
  • Hands-on with modern observability platforms (Prometheus, Grafana, OpenTelemetry, Datadog).
  • Incident management, capacity planning, and performance optimization experience.
  • Strong cloud and platform security knowledge (container/Kubernetes CI/CD security, vulnerability management).
  • Familiar with security frameworks (OWASP Top 10, IAM, secrets management, SSDLC).
  • Excellent collaboration, communication, and technical leadership.

Responsibilities

  • Define and improve CDI platform reliability using SLIs, SLOs, and error budgets.
  • Lead incident management, postmortems, and root-cause analysis across teams.
  • Develop comprehensive observability, logging, tracing, and alerting capabilities.
  • Build automation and self-service tools to reduce toil and boost productivity.
  • Design cloud-native infrastructure for scalable, cost-efficient platforms.
  • Drive IaC standards and deployment automation for consistency and agility.
  • Lead capacity planning and performance optimization initiatives.
  • Develop disaster recovery and business continuity strategies.
  • Partner with security to integrate SecOps practices into operations.
  • Mentor engineers and promote culture of reliability and security excellence.

Skills

12+ years SRE
Kubernetes
GCP
IaC
Linux
Networking
Distributed systems
Go
Python
Java
Shell
Incident management
Reliability
Observability
Cloud security
CI/CD security
Vulnerability management
OWASP Top 10
IAM
SSDLC
Security-by-design
Collaboration
Leadership

Education

Tools

Terraform
Prometheus
Grafana
OpenTelemetry
Datadog

Job description

Founded by fans, Crunchyroll delivers the art and culture of anime to a passionate community. We super-serve over 100 million anime and manga fans across 200+ countries and territories, and help them connect with the stories and characters they crave. Whether that experience is online or in-person, streaming video, theatrical, games, merchandise, events and more, it’s powered by the anime content we all love.

Join our team, and help us shape the future of anime!

About the role

We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the US and play a critical role in advancing the reliability, scalability, performance, and security of Crunchyroll's consumer-facing data platforms. As a senior technical leader, you will partner closely with Engineering, Data, Infrastructure, Product, and Security teams to design and operate resilient cloud-native systems that power critical business and customer experiences. You will drive initiatives across observability, incident management, automation, capacity planning, disaster recovery, and operational excellence while helping teams adopt modern SRE practices such as SLIs, SLOs, and error budgets.

The ideal candidate combines deep expertise in large-scale distributed systems with a strong sense of ownership, collaboration, and service leadership. You are passionate about building highly reliable platforms, eliminating operational toil through automation, and enabling engineering teams to move quickly and safely. In addition, you will champion SecOps best practices by driving vulnerability management, supporting penetration testing initiatives, improving security observability, strengthening cloud and Kubernetes security controls, and ensuring operational readiness for emerging threats. This is a unique opportunity to shape reliability and security engineering practices across CDI while helping build a world-class data and insights ecosystem that enables informed decision-making throughout Crunchyroll.

Core Areas of Responsibility
  • Reliability Engineering: Define, measure, and continuously improve the reliability, availability, and performance of CDI platforms through SLIs, SLOs, and error budgets.
  • Operational Excellence: Establish and drive best practices for incident management, root cause analysis, postmortems, and service ownership across engineering teams.
  • Observability & Monitoring: Build and evolve comprehensive monitoring, logging, tracing, and alerting capabilities to enable proactive issue detection and rapid resolution.
  • Automation: Identify operational inefficiencies and develop automation, self-service capabilities, and self-healing mechanisms to improve engineering productivity.
  • Platform Scalability: Design and optimize cloud-native infrastructure and services to support growing business demands while maintaining performance and cost efficiency.
  • Infrastructure Engineering: Drive Infrastructure as Code (IaC), platform standardization, and deployment automation to improve consistency, reliability, and operational agility.
  • Capacity Planning & Performance: Lead capacity planning and performance optimization initiatives to ensure platforms can scale predictably and efficiently.
  • Disaster Recovery & Resilience: Develop and regularly validate disaster recovery, backup, and business continuity strategies to ensure platform resiliency.
  • Security Operations (SecOps): Partner with Crunchyroll's security team to integrate security controls, operational risk management, and security best practices into platform operations and engineering workflows.
  • Vulnerability Management: Own the triage and remediation of identified vulnerabilities across infrastructure, platform, container, and application security vulnerabilities through established Crunchyroll vulnerability management processes.
  • Penetration Testing & Security Remediation: Support penetration test scoping activities by providing technical context on CDI platforms. Own the triage, prioritization, and remediation of resulting findings to drive timely resolution and strengthen platform security posture.
  • Cloud & Kubernetes Security: Implement and maintain secure cloud, container, and Kubernetes environments following least-privilege, defense-in-depth, and Zero Trust principles.
  • Cross-Functional Leadership: Collaborate with Engineering, Data, Product, Infrastructure, and Security teams to drive reliability, scalability, and security initiatives across CDI.
  • Mentorship & Engineering Excellence: Mentor engineers and champion a culture of operational excellence, reliability, ownership, continuous improvement, and security awareness.
About You

We get excited about candidates like you, because…

  • 12+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Engineering, or related disciplines, with a proven track record of operating and scaling production-critical systems.
  • Deep expertise in Kubernetes and GCP, including the design, deployment, and operation of highly available, cloud-native platforms at scale.
  • Strong Infrastructure as Code (IaC) experience, preferably with Terraform, and a commitment to automation, standardization, and operational efficiency.
  • Solid foundation in Linux systems administration, networking, and distributed systems, with the ability to troubleshoot complex production issues across multiple layers of the technology stack.
  • Proficiency in one or more programming and scripting languages, such as Go, Python, Java, or Shell, with a focus on automation and platform engineering.
  • Hands-on experience with modern observability platforms and practices, including Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent monitoring and telemetry solutions.
  • Demonstrated expertise in incident management, service reliability, capacity planning, performance optimization, and operational excellence, including the implementation of SLIs, SLOs, and error budgets.
  • Strong understanding of cloud and platform security, including container security, Kubernetes security, CI/CD security, vulnerability management, and secure infrastructure operations.
  • Good to have knowledge of security frameworks and best practices, including OWASP Top 10, Identity and Access Management (IAM), secrets management, Secure Software Development Lifecycle (SSDLC), and security-by-design principles.
  • Excellent collaboration, communication, and technical leadership skills, with experience influencing architectural decisions, driving cross-functional initiatives, and mentoring engineers in reliability and operational best practices.
About the Team

The Center for Data and Insights (CDI) is a service-oriented, horizontal organization uniquely positioned within the company to serve as the trusted, unbiased source of timely, data-driven insights for Crunchyroll. Our vision is to inspire, support, and guide our stakeholders to be data-aware and build the systems of intelligence to discover insights and act on them. We have built a highly functional organization that truly believes in being a responsive partner, with the utmost curiosity, unwavering accountability and the courage to lead with actions.

Why you will love working at Crunchyroll

In addition to getting to work with fun, passionate and inspired colleagues, you will also enjoy the following benefits and perks:

  • Receive a great compensation package including salary plus performance bonus earning potential, paid annually.
  • Flexible time off policies allowing you to take the time you need to be your whole self.
  • Generous medical, dental, vision, STD, LTD, and life insurance
  • Health Saving Account HSA program
  • Health care and dependent care FSA
  • 401(k) plan, with employer match
  • Support program for new parents
  • Pet insurance and some of our offices are pet friendly!

#LifeAtCrunchyroll #LI-Hybrid

The Pay Range for this position is listed. Actual pay will vary based on factors including, but not limited to location, experience, and performance. The range listed is just one component of Crunchyroll’s Total Rewards offerings for employees. Other rewards may include performance bonuses, employer matched retirement savings, time-off programs, and progressive health benefits and perks.

$233,400 - $291,800 USD

About our Values

We want to be everything for someone rather than something for everyone and we do this by living and modeling our values in all that we do.

Courage. We believe that when we overcome fear, we enable our best selves.

Curiosity. We are curious, which is the gateway to empathy, inclusion, and understanding.

  • Kaizen. We have a growth mindset committed to constant forward progress.

Service. We serve our community with humility, enabling joy and belonging for others.

Our commitment to diversity and inclusion

Our mission of helping people belong reflects our commitment to diversity & inclusion. It's just the way we do business.

We are an equal opportunity employer and value diversity at Crunchyroll. Pursuant to applicable law, we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Crunchyroll, LLC is an independently operated joint venture between US-based Sony Pictures Entertainment, and Japan's Aniplex, a subsidiary of Sony Music Entertainment (Japan) Inc., both subsidiaries of Tokyo-based Sony Group Corporation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

Ellation, Inc. • Los Angeles (CA)

On-site
USD 211,000 - 263,000
Salary plus performance bonus
Flexible time off
Medical, dental, vision insurance
+2
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Crunchyroll, LLC • Los Angeles (CA)

On-site
USD 211,000 - 263,000
Salary plus performance bonus
Flexible time off
Medical, dental, vision insurance
+2
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Crunchyroll • San Francisco (CA)

On-site
USD 233,000 - 292,000
Performance bonus
Flexible time off
Medical insurance
+5
Software Engineer III, Media Delivery
Software Engineer III, Media Delivery

Crunchyroll, LLC • San Francisco (CA)

Hybrid
USD 169,000 - 211,000
Salary plus performance bonus
Flexible time off
Medical, dental, vision, STD, LTD, and
+5
Data Engineer III
Data Engineer III

Crunchyroll • San Francisco (CA)

On-site
USD 164,100 - 205,100
Performance bonus potential
Flexible time off policies
Health, dental, and vision insurance
+3
Senior Software Engineer, Infrastructure Engineering
Senior Software Engineer, Infrastructure Engineering

Crunchyroll • Los Angeles (CA)

On-site
USD 183,000 - 229,000
Salary plus performance bonus
Flexible time off
Medical, dental, vision
+5
Senior Database Engineer
Senior Database Engineer

Crunchyroll • San Francisco (CA)

On-site
USD 203,000 - 254,000
Salary + bonus
Flexible time off
Medical/Dental/Vision
+7
Software Engineer III, Media Delivery
Software Engineer III, Media Delivery

Crunchyroll • San Francisco (CA)

On-site
USD 169,000 - 211,000
Performance bonus
Flexible time off
Medical, dental, vision insurance
+3
Senior Software Engineer, Infrastructure Engineering
Senior Software Engineer, Infrastructure Engineering

Crunchyroll, LLC • Los Angeles (CA)

On-site
USD 183,000 - 229,000
Salary plus performance bonus
Flexible time off
Medical, dental, vision, STD, LTD, and
+5
Senior Program Manager - Customer Experience
Senior Program Manager - Customer Experience

Crunchyroll, LLC • Los Angeles (CA)

Hybrid
USD 144,000 - 160,000
Salary plus performance bonus
Flexible time off
Medical, dental, vision insurance
+2