Lead Site Reliability Engineer

Disney

Bengaluru

On-site

INR 2,500,000 - 5,000,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Disney Experience platforms require a seasoned Lead Site Reliability Engineer to drive reliability, scalability, and performance worldwide. You will guide other engineers, set technical direction for complex systems, and ensure highly available, secure experiences for guests across parks, resorts, and digital touchpoints.

You will mentor teams, own SLIs/SLOs/SLAs, and lead incident postmortems while advancing DevOps, automation, and FinOps practices in a fast-paced, global environment.

Qualifications

  • 7+ years in SRE/Systems Eng or similar, with large-scale cloud systems.
  • Degree or equivalent in a technical field.
  • Hands-on with AWS/Azure/GCP and container orchestration.
  • IaC and CI/CD tooling experience.
  • Fluency in scripting and programming languages.
  • Observability tools and incident response expertise.

Responsibilities

  • Architect, design, and build scalable, secure infrastructure and platforms.
  • Lead CI/CD pipelines, automation, and telemetry for monitoring and incidents.
  • Own reliability strategies, SLIs/SLOs/SLAs, and uptime improvements.
  • Drive root-cause analyses and remediation plans for incidents.
  • Develop automation and IaC to standardize deployments.
  • Collaborate across product, engineering, security, and ops teams.
  • Mentor engineers and guide team projects toward reliable outcomes.
  • Apply FinOps principles to balance cost, performance, and reliability.
  • Foster an inclusive, innovative engineering culture.

Skills

SRE
Cloud architecture
CI/CD
Infrastructure as Code
Observability
Networking
Unix/Linux
Programming scripting

Education

Bachelor's in CS/IS/Engineering or equivalent

Tools

Docker
Kubernetes
ECS
AKS
GKE
Terraform
CloudFormation
Ansible
Chef
GitHub
GitLab
Jenkins
AWS CodeBuild
Azure DevOps

Job description

Lead Site Reliability Engineer

Req ID: 10159535

Job Description

The Lead Site Reliability Engineer is a seasoned subject matter expert who drives reliability, scalability, and performance for critical Disney Experiences platforms that power immersive guest interactions across theme parks, resorts, cruise, vacation, travel, retail, and consumer experiences. In this lead role, you will guide other engineers, set technical direction for complex systems, and ensure our digital and physical experiences remain highly available, secure, and resilient for guests around the world.

Responsibilities
  • Architect, design, and build scalable, maintainable, and secure infrastructure and platforms, including cloud-native and container-based solutions, to support mission-critical commerce and guest-facing applications.
  • Lead the evolution of DevOps and SRE practices by consulting on, designing, and supporting CI/CD pipelines, automating infrastructure and operations, and creating telemetry and observability for monitoring and incident response.
  • Serve as the SRE subject matter expert and technical lead for assigned products and platforms, owning reliability strategies, defining SLIs/SLOs/SLAs, and driving continuous improvement in uptime and performance.
  • Identify root causes of operational issues in large-scale distributed systems, lead major incident response, and deliver clear retrospectives and remediation plans that reduce future risk and operational toil.
  • Develop, maintain, and enhance automation, scripts, and Infrastructure as Code to standardize deployments, improve reliability, and support complex, non-standard environments without relying solely on runbooks.
  • Collaborate with product, engineering, security, and operations teams to plan capacity, monitoring, configuration, security, metrics, reporting, recovery, and migration strategies for new initiatives and events impacting supported platforms.
  • Mentor, train, and guide other engineers by providing continuous coaching, feedback, and technical direction, holding self and others accountable to commitments and aligning team work with organizational goals.
  • Plan and coordinate team efforts and platform-oriented projects with moderate complexity and risk, breaking down organizational goals into clear outcomes and negotiating solutions to complex reliability challenges.
  • Apply FinOps and cost-optimization principles to cloud environments, implementing governance, tagging, rightsizing, and usage analysis to balance reliability, performance, and cost efficiency.
  • Champion a diverse, inclusive, team-oriented culture that encourages innovation, creative problem solving, and service-minded collaboration, ensuring every voice is heard and Disney values are experienced daily.
Required Qualifications
  • Minimum 7 years of related work experience in Site Reliability Engineering, Systems Engineering, or software development, with a focus on large-scale, distributed, and cloud-based systems.
  • Bachelor's Degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent work experience.
  • Extensive hands-on experience with cloud hosting services (AWS, Azure, Google Cloud) and modern cloud architectures, including containers and orchestration platforms such as Docker, Kubernetes, ECS, AKS, and GKE.
  • Proficiency in Infrastructure as Code and configuration management tools (e.g., Terraform, CloudFormation, Ansible, Chef) and CI/CD pipelines using tools such as GitHub, GitLab, Jenkins, AWS CodeBuild, or Azure DevOps.
  • Fluency in core scripting and programming languages (e.g., Python, NodeJS, Golang, Bash, Perl, Ruby, Java) and strong UNIX/Linux administration, troubleshooting, and security skills.
  • Applied expertise in observability and monitoring, including defining and implementing SLIs, SLOs, SLAs and using major APM and logging tools (e.g., AppDynamics, New Relic, ELK stack, Datadog, Splunk, New Relic).
  • Strong knowledge of networking and distributed systems, including HTTP, TCP/IP, DNS, TLS, SSH, VPCs, gateways, firewalls, and microservices architectures.
  • Experience with databases and data platforms such as MySQL, MongoDB, DynamoDB, Redis, and data solutions like Snowflake or Tableau, including ELT processes for data-driven decision m
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

The Walt Disney Company (France) • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JobCubby • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Mgr, Site Reliability Engineering
Mgr, Site Reliability Engineering

Disney • Bengaluru

On-site
INR 4,000,000 - 7,500,000
Mgr, Site Reliability Engineering
Mgr, Site Reliability Engineering

Disney Cruise Line - The Walt Disney Company • Bengaluru

On-site
Confidential
Sr Site Reliability Engineer
Sr Site Reliability Engineer

The Walt Disney Company (France) • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Dir, Software Engineering
Dir, Software Engineering

Disney • Bengaluru

On-site
INR 4,200,000 - 6,000,000
Site Reliability Engineer
Site Reliability Engineer

The Walt Disney Company (France) • Bengaluru

On-site
INR 8,507,000 - 12,287,000
Mgr, Software Engineering
Mgr, Software Engineering

Disney Experiences • Bengaluru

On-site
INR 4,500,000 - 7,500,000
Sr Database Administrator
Sr Database Administrator

The Walt Disney Company (France) • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Site Reliability Engineering Lead_Truist
Site Reliability Engineering Lead_Truist

Infosys • Bengaluru

On-site
INR 4,000,000 - 7,000,000