Senior Site Reliability Engineer

Ellation, Inc.

Hyderabad

Hybrid

INR 3,500,000 - 6,000,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Best-in-class medical insurance
24/7 EAP counseling
Free Crunchyroll premium
Professional development
Parental leave up to 26 weeks
Hybrid work schedule
Paid time off
Flex time off
Yasumi days
Half-day Fridays (summer)
Winter break

Job summary

Crunchyroll is hiring a Staff Site Reliability Engineer to join the Center for Data & Insights (CDI) in India. The role focuses on reliability, scalability, security, and observability for data platforms serving global audiences.

You will partner with Engineering, Data, Infrastructure and Security to operate resilient cloud-native systems and drive initiatives in SLIs, SLOs, and automation. Ideal candidates bring 8+ years in SRE/Platform Engineering, strong Kubernetes and GCP experience, and

Qualifications

  • 8+ years of experience in SRE/Platform/Infra engineering or related
  • Strong experience with Kubernetes and GCP at scale
  • Excellent IaC experience, preferably Terraform
  • Solid Linux and networking fundamentals
  • Programming/scripting in Go, Python, Java, or Shell
  • Proficient in modern observability platforms (Prometheus, Grafana, OpenTelemetry, Datadog)
  • Experience with incident management, SLIs/SLOs, and operational excellence
  • Platform security experience including container and Kubernetes security
  • Familiarity with security best practices and IAM, secrets management, SDLC
  • Strong collaboration and problem-solving across teams

Responsibilities

  • Define and improve reliability, availability, and performance of CDI platforms using SLIs, SLOs, and error budgets
  • Lead incident management, postmortems, RCA, and service ownership across engineering teams
  • Build and evolve monitoring, logging, tracing, and alerting capabilities for proactive issue detection
  • Develop automation and self-service capabilities to reduce toil
  • Design and optimize cloud-native infrastructure for scale, performance, and cost efficiency
  • Implement IaC and deployment automation for consistency and agility
  • Participate in capacity planning and performance optimization
  • Develop and validate disaster recovery and business continuity strategies
  • Collaborate with Security to integrate SecOps practices into platform operations
  • Triage and remediate vulnerabilities across infrastructure and containers
  • Support penetration testing scoping with technical context and remediation prioritization
  • Maintain secure cloud and Kubernetes environments following least-privilege and Zero Trust principles
  • Collaborate across Engineering, Data, Product, Infrastructure, and Security teams

Skills

Kubernetes
GCP
Terraform
IaC
Linux
Networking
Go/Python/Java/Shell
Observability
Prometheus
Grafana/OpenTelemetry/Datadog
SLIs/SLOs/Incident mgmt
SecOps/Security
Cloud & Kubernetes Security
Collaboration & Communication

Tools

Terraform
Datadog
Prometheus
OpenTelemetry
Grafana

Job description

Founded by fans, Crunchyroll delivers the art and culture of anime to a passionate community. We super-serve over 100 million anime and manga fans across 200+ countries and territories, and help them connect with the stories and characters they crave. Whether that experience is online or in-person, streaming video, theatrical, games, merchandise, events and more, it’s powered by the anime content we all love.

Join our team, and help us shape the future of anime!

About the role

We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in India and play a critical role in advancing the reliability, scalability, performance, and security of Crunchyroll's consumer-facing data platforms. As a senior technical leader, you will partner closely with Engineering, Data, Infrastructure, Product, and Security teams to design and operate resilient cloud-native systems that power critical business and customer experiences. You will drive initiatives across observability, incident management, automation, capacity planning, disaster recovery, and operational excellence while helping teams adopt modern SRE practices such as SLIs, SLOs, and error budgets.

The ideal candidate combines deep expertise in large-scale distributed systems with a strong sense of ownership, collaboration, and service leadership. You are passionate about building highly reliable platforms, eliminating operational toil through automation, and enabling engineering teams to move quickly and safely. In addition, you will champion SecOps best practices by driving vulnerability management, supporting penetration testing initiatives, improving security observability, strengthening cloud and Kubernetes security controls, and ensuring operational readiness for emerging threats. This is a unique opportunity to shape reliability and security engineering practices across CDI while helping build a world-class data and insights ecosystem that enables informed decision-making throughout Crunchyroll.

Core Areas of Responsibility
  • Reliability Engineering: Define, measure, and continuously improve the reliability, availability, and performance of CDI platforms through SLIs, SLOs, and error budgets.
  • Operational Excellence: Establish and support best practices for incident management, root cause analysis, postmortems, and service ownership across engineering teams.
  • Observability & Monitoring: Build and evolve comprehensive monitoring, logging, tracing, and alerting capabilities to enable proactive issue detection and rapid resolution.
  • Automation: Develop automation, self-service capabilities, and self-healing mechanisms to improve engineering productivity.
  • Platform Scalability: Design and optimize cloud-native infrastructure and services to support growing business demands while maintaining performance and cost efficiency.
  • Infrastructure Engineering: Implement Infrastructure as Code (IaC), platform standardization, and deployment automation to improve consistency, reliability, and operational agility.
  • Capacity Planning & Performance: Participate in capacity planning and performance optimization initiatives to ensure platforms can scale predictably and efficiently.
  • Disaster Recovery & Resilience: Develop and regularly validate disaster recovery, backup, and business continuity strategies to ensure platform resiliency.
  • Security Operations (SecOps): Partner with Crunchyroll's security team to integrate security controls, operational risk management, and security best practices into platform operations and engineering workflows.
  • Vulnerability Management: Support the triage and remediation of identified vulnerabilities across infrastructure, platform, container, and application security vulnerabilities through established Crunchyroll vulnerability management processes.
  • Penetration Testing & Security Remediation: Support penetration test scoping activities by providing technical context on CDI platforms. Own the triage, prioritization, and remediation of resulting findings to drive timely resolution and strengthen platform security posture.
  • Cloud & Kubernetes Security: Implement and maintain secure cloud, container, and Kubernetes environments following least-privilege, defense-in-depth, and Zero Trust principles.
  • Cross-Functional Collaboration: Collaborate with Engineering, Data, Product, Infrastructure, and Security teams to drive reliability, scalability, and security initiatives across CDI.
About You

We get excited about candidates like you, because…

  • 8+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Engineering, or related disciplines, with a proven track record of operating and scaling production-critical systems.
  • Strong experience with Kubernetes and GCP, including deploying, operating, and troubleshooting cloud-native applications and services at scale.
  • Excellent Infrastructure as Code (IaC) Experience in implementing IaC solutions, preferably using Terraform, to improve automation, consistency, and operational efficiency.
  • Systems & Networking Fundamentals: Solid understanding of Linux systems administration, networking fundamentals, and distributed systems concepts
  • Programming & Automation: Proficiency in one or more programming or scripting languages such as Go, Python, Java, or Shell
  • Observability & Monitoring: Experience with modern observability platforms, including Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent monitoring and telemetry solutions.
  • Service Reliability & Operations: Working knowledge of incident management, service reliability practices, capacity planning, performance optimization, and operational excellence, including the use of SLIs and SLOs.
  • Platform Security: Experience supporting cloud and platform security initiatives, including container security, Kubernetes security, CI/CD security, vulnerability remediation, and secure infrastructure operations.
  • Security Best Practices: Familiarity with industry-standard security frameworks and practices, including OWASP Top 10, Identity and Access Management (IAM), secrets management, SSDLC, and security-by-design principles.
  • Collaboration & Communication: Strong communication, collaboration, and problem-solving skills, with the ability to work effectively across teams and contribute to reliability and operational excellence initiatives.
About the Team

The Center for Data and Insights (CDI) is a service-oriented, horizontal organization uniquely positioned within the company to serve as the trusted, unbiased source of timely, data-driven insights for Crunchyroll. Our vision is to inspire, support, and guide our stakeholders to be data-aware and build the systems of intelligence to discover insights and act on them. We have built a highly functional organization that truly believes in being a responsive partner, with the utmost curiosity, unwavering accountability, and the courage to lead with actions.

Why you will love working at Crunchyroll

In addition to getting to work with fun, passionate and inspired colleagues, you will also enjoy the following benefits and perks:

  • Best-in class medical, dental, and vision private insurance healthcare coverage
  • Access to counseling & mental health sessions 24/7 through our Employee Assistance Program (EAP)
  • Free premium access to Crunchyroll
  • Professional Development
  • Company's Paid Parental Leave
  • up to 26 weeks for birthing parents
  • up to 12 weeks for non-birthing parents
  • Hybrid Work Schedule
  • Paid Time Off
  • Flex Time Off
  • 5 Yasumi Days
  • Half-Day Fridays during the summer
  • Winter Break

#LifeAtCrunchyroll ((select from the following job modalities for this role: #LI-Hybrid #LI-remote #LI-onsite))

About our Values

We want to be everything for someone rather than something for everyone and we do this by living and modeling our values in all that we do. We value

Courage. We believe that when we overcome fear, we enable our best selves.

Curiosity. We are curious, which is the gateway to empathy, inclusion, and understanding.

  • Kaizen. We have a growth mindset committed to constant forward progress.

Service. We serve our community with humility, enabling joy and belonging for others.

Our commitment to diversity and inclusion

Our mission of helping people belong reflects our commitment to diversity & inclusion. It's just the way we do business.

We are an equal opportunity employer and value diversity at Crunchyroll. Pursuant to applicable law, we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Crunchyroll, LLC is an independently operated joint venture between US-based Sony Pictures Entertainment, and Japan's Aniplex, a subsidiary of Sony Music Entertainment (Japan) Inc., both subsidiaries of Tokyo-based Sony Group Corporation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Crunchyroll • Hyderabad

On-site
INR 3,000,000 - 5,500,000
Best-in-class medical, dental, and eye
Hybrid work schedule
Paid time off
Staff Software Engineer
Staff Software Engineer

Crunchyroll • Hyderabad

On-site
INR 4,000,000 - 6,400,000
Medical, dental & vision insurance
Hybrid work schedule
Professional development
+2
Senior Data Analyst
Senior Data Analyst

Crunchyroll • Hyderabad

On-site
INR 1,800,000 - 2,400,000
Medical insurance
Dental and vision coverage
Hybrid work schedule
+2
Senior Software Engineer -Backend/Full Stack
Senior Software Engineer -Backend/Full Stack

Crunchyroll • Hyderabad

On-site
INR 2,400,000 - 4,200,000
Senior Data Analyst Crunchyroll
Senior Data Analyst Crunchyroll

Earn Modes • Hyderabad

Hybrid
INR 2,600,000 - 3,400,000
Senior Database Operations Engineer
Senior Database Operations Engineer

Crunchyroll • Hyderabad

Hybrid
INR 900,000 - 1,500,000
Medical, dental, and vision insurance
Employee Assistance Program (EAP)
Free Crunchyroll access
+6
Staff Software Engineer, AI/ML
Staff Software Engineer, AI/ML

Crunchyroll, LLC • Hyderabad

On-site
INR 4,500,000 - 6,500,000
Senior MLOps Engineer
Senior MLOps Engineer

Crunchyroll • Hyderabad

Hybrid
INR 1,500,000 - 2,500,000
Best-in-class medical, dental, and vision private insurance
Free premium access to Crunchyroll
Paid parental leave up to 26 weeks
+2
Delivery Operations Coordinator
Delivery Operations Coordinator

crunchyroll • Hyderabad

Hybrid
INR 550,000 - 900,000
Senior Software Engineer, Devops/SRE
Senior Software Engineer, Devops/SRE

Roku, Inc. • Bengaluru

Hybrid
INR 1,500,000 - 2,000,000
Mental health and financial wellness support
Healthcare benefits
401(k)/pension options