Site Reliability Engineer II

The Walt Disney Company

New York (NY)

On-site

USD 123,000 - 165,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The Walt Disney Company is seeking a Site Reliability Engineer II based in New York City. This role involves enhancing system reliability, performance optimization, and building automation. The ideal candidate possesses extensive experience in Site Reliability Engineering and is proficient in cloud platforms like AWS.

Key responsibilities include collaborating with teams to implement SRE practices, supporting CI/CD pipelines, and contributing to incident response efforts. Competitive compensation ranges from $123,000 to $165,000 annually, depending on experience.

Qualifications

  • 3+ years of experience in Site Reliability Engineering, DevOps, or related discipline.
  • Hands-on experience with cloud platforms - AWS preferred.
  • Familiarity with Infrastructure-as-Code (Terraform, CloudFormation).

Responsibilities

  • Contribute to the design and improvement of systems for reliability and performance.
  • Build automation for deployment and operational workflows.
  • Collaborate with engineering teams to implement SRE best practices.
  • Participate in incident response and root cause analysis.

Skills

Site Reliability Engineering
Cloud platforms (AWS preferred)
Python
Go
JavaScript
CI/CD systems
Containerization (Docker, Kubernetes)

Education

Bachelor's degree in computer science or related field

Tools

Terraform
Prometheus
Grafana

Job description

Site Reliability Engineer II

Job ID 10143234
Location New York, New York, United States
Business Disney Entertainment and ESPN Product & Technology
Date posted May 19, 2026

Job Description

The Streaming SRE squad drives improvements in performance, resiliency, and operational excellence. We take a consultative approach to reliability engineering—partnering with a variety of cross‑functional teams to provide guidance, automation, education, and best practices that elevate the reliability and scalability of services that support our products and brands.

We are seeking a Site Reliability Engineer who will contribute to the stability and scalability of critical systems by building automation, improving operational workflows, enhancing observability, and participating in incident response. The ideal candidate has a strong understanding of distributed system fundamentals, cloud‑native resources and operations, and performance optimization. This role requires a collaborative mindset and the ability to work closely with engineering teams to implement SRE principles across the organization.

Fostering innovation is a critical component to success here at Disney Entertainment and ESPN Product & Technology. Therefore, the ideal candidate will also need to be highly adaptable to changes and be able to pivot when required.

Responsibilities
  • Contribute to the design, implementation, and improvement of systems to enhance reliability, scalability, and performance.
  • Build and maintain automation for deployment, monitoring, alerting, and operational workflows.
  • Collaborate with software engineering teams to implement SRE best practices, including SLIs, SLOs, error budgets, and automated remediation.
  • Support CI/CD pipelines and participate in optimizing the software delivery lifecycle.
  • Develop tools, dashboards, and instrumentation to improve observability across metrics, logs, and distributed tracing.
  • Participate in incident response, root cause analysis (RCA), and corrective actions to prevent recurrence.
  • Assist in capacity planning, performance tuning, and scaling strategies for distributed systems.
  • Maintain and improve Infrastructure‑as‑Code (IaC) definitions and cloud environment configurations.
  • Contribute to documentation, runbooks, architectural diagrams, and operational standards.
  • Collaborate with cross‑functional teams to identify reliability risks and recommend improvements.
  • Participate in incident‑based escalations and rotations to support high‑availability production systems.
  • Continuously evaluate system architecture, tools, and practices to drive operational excellence and efficiency.
Basic Qualifications
  • Bachelor's degree in computer science, engineering, or related field (or equivalent experience).
  • 3+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related discipline.
  • Hands‑on experience with cloud platforms – AWS (preferred), GCP, Azure.
  • Proficiency in Python, Go, JavaScript, Bash, or equivalent scripting languages.
  • Working knowledge of Linux or Unix‑based systems.
  • Experience with CI/CD systems (e.g., GitHub Actions, GitLab CI, Jenkins).
  • Familiarity with Infrastructure‑as‑Code (Terraform, CloudFormation, etc.).
  • Experience with containerization technologies such as Docker and Kubernetes.
  • Understanding networking fundamentals, distributed systems, and system design basics.
  • Strong analytical and troubleshooting skills, including the ability to diagnose complex system issues.
  • Ability to work both independently and collaboratively.
  • Strong communication skills and the ability to collaborate effectively with cross‑functional teams.
Preferred Qualifications
  • Hands‑on experience with observability stacks (Prometheus, Grafana, ELK/EFK, Datadog, Splunk, New Relic).
  • Exposure to GitOps tooling (Argo CD, Flux).
  • Experience contributing to SLO/SLI frameworks and implementing error budgets.
  • Knowledge of service mesh architectures (Istio, Linkerd).
  • Familiarity with performance testing and load testing tools.
  • Experience with message queues, event‑driven systems, or distributed data platforms.
  • Cloud or DevOps‑related certifications (AWS Associate/Specialty, GCP Professional, Kubernetes CKA/CKS).
  • Experience working in large‑scale enterprise environments or with distributed global teams.
  • Experience using modern AI‑assisted development tools (e.g., Copilot, Cursor) to improve code quality, accelerate development, and enhance documentation.
  • Understanding foundational AI/ML concepts, familiarity with cloud‑native AI services such as model hosting, and/or ability to use AI tools to automate cloud operations tasks.
Compensation

The hiring range for this position in New York City is $123,000 - $165,000. The base pay actually offered will take into account internal equity and may vary depending on the candidate’s geographic region, job‑related knowledge, skills, and experience among other factors. A bonus and/or long‑term incentive units may be provided as part of the compensation package, in addition to the full range of medical, financial, and/or other benefits, dependent on the level and position offered.

Equal Opportunity Employer Statement

Disney Entertainment & Sports LLC is an equal opportunity employer. Applicants will receive consideration for employment without regard to race, religion, color, sex, sexual orientation, gender, gender identity, gender expression, national origin, ancestry, age, marital status, military or veteran status, medical condition, genetic information or disability, or any other basis prohibited by federal, state or local law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

5014 Disney Entertainment & Sports LLC • New York (NY)

On-site
USD 123,000 - 165,000
Streaming SRE: Reliability & Automation Engineer
Streaming SRE: Reliability & Automation Engineer

The Walt Disney Company • New York (NY)

On-site
USD 123,000 - 165,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

5131 Hulu Enterprises, LLC • United States

On-site
USD 140,000 - 180,000
Software Engineer - DevOps Mobile
Software Engineer - DevOps Mobile

The Walt Disney Company • New York (NY)

On-site
USD 123,000 - 165,000
Lead Software Engineer
Lead Software Engineer

The Walt Disney Company • New York (NY)

On-site
USD 163,000 - 219,000
Software Engineer - Java
Software Engineer - Java

The Walt Disney Company • New York (NY)

On-site
USD 123,000 - 165,000
Medical benefits
Bonus opportunities
Long-term incentives
Sr Software Engineer
Sr Software Engineer

The Walt Disney Company • Santa Monica (CA)

On-site
USD 138,900 - 186,200
Site Reliability Engineer II: Scale, Automation & Observability
Site Reliability Engineer II: Scale, Automation & Observability

5014 Disney Entertainment & Sports LLC • New York (NY)

On-site
USD 123,000 - 165,000
Principal Software Engineer - Observability
Principal Software Engineer - Observability

The Walt Disney Company • Glendale (CA), New York (NY)

On-site
USD 184,000 - 259,000
Product Software Engineer I
Product Software Engineer I

The Walt Disney Company • San Francisco (CA), Seattle (WA), New York (NY)

On-site
USD 102,000 - 143,000