Staff Site Reliability Engineer Playout

NBCUniversal

Stamford (CT)

On-site

USD 145,000 - 175,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NBCUniversal is seeking a Staff Site Reliability Engineer (Playout) to drive reliability for live linear playout systems across NBCUniversal channels. The role leads a team of SREs, defines service levels, and shapes resiliency strategy in a fast-paced environment.

You will own incident response, capacity planning, automation, and monitoring with Grafana dashboards, ensuring predictable performance and availability for cloud-based playout services.

Qualifications

  • Bachelor's degree in computer science or related field.
  • 8+ years hands-on engineering in broadcast playout environments (e.g., Snell, Harris, Imagine, Amagi).
  • Willingness to on-call 24/7 for escalations.
  • Experience with monitoring/logging tools (Splunk, Grafana).
  • Experience with streaming protocols and codecs (TS, HEVC, H.264, HLS, CMAF, SCTE-35/224, ESAM, SRT/RIST).
  • Experience with IP networking and cloud-based networks (AWS).
  • Docker & Kubernetes experience; excellent communicator; automation-first mindset.

Responsibilities

  • Team lead for SRE engineers on the playout engineering team.
  • Define and manage reliability targets (SLIs/SLOs) and readiness criteria.
  • Drive incident response; lead major incident management and post-incident reviews.
  • Collaborate with engineering, product, and operations to improve reliability via capacity planning and tuning.
  • Provide runbooks and architecture support for reviews and planning.
  • Lead automation efforts to reduce toil and improve MTTR.
  • Provide L1/L2 on-call support for playout infrastructure as needed.
  • Create monitoring dashboards and alerts using Grafana, Slack, and ServiceNow.

Skills

SRE
Broadcast playout
AWS
Docker
Kubernetes
Grafana
Splunk
On-call 24/7
SLIs/SLOs
Automation

Education

Bachelor's degree in computer science

Tools

Snell
Harris
Imagine
Amagi

Job description

Staff Site Reliability Engineer, Playout
  • Full-time
  • Business Segment: Operations & Technology
  • Compensation: USD 145,000 - USD 175,000 - yearly

NBCUniversal is one of the world's leading media and entertainment companies. We create world-class content, which we distribute across our portfolio of film, television, and streaming, and bring to life through our global theme park destinations, consumer products, and experiences. We own and operate leading entertainment and news brands, including NBC, NBC News, NBC Sports, Telemundo, NBC Local Stations, Bravo, and Peacock, our premium ad-supported streaming service. We produce and distribute premier filmed entertainment and programming through our powerhouse film and television studios, including Universal Pictures, DreamWorks Animation, and Focus Features, and the four global television studios under the Universal Studio Group banner, and operate industry-leading theme parks and experiences around the world through Universal Destinations & Experiences, including Universal Orlando Resort, home to Universal Epic Universe, and Universal Studios Hollywood. NBCUniversal is a subsidiary of Comcast Corporation. Visit www.nbcuniversal.com for more information.

Our impact is rooted in improving the communities where our employees, customers, and audiences live and work. We have a rich tradition of giving back and ensuring our employees have the opportunity to serve their communities. We champion an inclusive culture and strive to attract and develop a talented workforce to create and deliver a wide range of content reflecting our world.

NBCUniversal Operations & Technology is looking for a Staff SRE, Playout Engineering to provide technical leadership to a team of Site Reliability Engineers. This team drives reliability, observability, and operational excellence for cloud-based master control playout systems supporting all NBCUniversal live linear channels – NBC, Telemundo, Peacock Virtual Channels, etc.

In this position, you will shape the reliability strategy for live linear playout—defining service levels, improving resiliency, and strengthening monitoring and incident response—so the platform can meet evolving business needs with predictable performance and availability.

This role requires the ability to operate in a fast-paced environment. For systems in production, you will lead an on-call team and drive L1 and L2 troubleshooting, incident management and continuous improvement to maintain reliable distribution.

Responsibilities

  • Team lead for SRE engineers on playout Engineering team
  • Define and manage reliability targets (SLIs/SLOs) and operational readiness criteria for playout services
  • Drive incident response: establish on-call practices, lead major incident management, and ensure post-incident reviews result in measurable improvements
  • Partner with engineering, product, and operations teams to improve reliability through capacity planning, performance tuning, and resilience testing
  • Provide high-level conceptual drawings and operational runbooks to support architecture reviews, support readiness, and project planning
  • Leadership in driving automation to reduce toil and improve reliability and mean time to recovery (MTTR)
  • L1 & L2 support to maintain playout infrastructure/services for NBCUniversal including providing after hours on-call support select week
  • Leadership in creating monitoring dashboards (Grafana) and proper alerts(teams/slack/ServiceNow)

Qualifications/Requirements

  • Bachelor's degree in computer science or related degree / experience
  • Eight years’ hands-on-keyboard Engineering experience working with broadcast automation playout environments e.g. Snell, Harris, Imagine, Amagi
  • Requires on-call 24/7 availability for escalations
  • Experience with monitoring/logging tools e.g. Splunk and Grafana
  • Experience with streaming protocols and codecs (e.g. TS, HEVC, H.264, HLS, CMAF, SCTE-35, SCTE-224, ESAM, SRT/RIST)
  • Experience with IP networking and interfacing with cloud-based networks
  • Experience with containerization (Docker & Kubernetes)
  • Excellent communicator and able to clearly articulate complex issues and technologies
  • Expert with broadcast playout systems (master control) technologies
  • Expert with public cloud environments using AWS services
  • Comfortable working in a fast-paced agile environment. Requirements change quickly and our team needs to constantly adapt to meet objectives
  • An automate-first and automate everything attitude

Desired Characteristics

  • Experience with cloud native playout vendor solutions (Amagi, Evertz, GrassValley, Harmonic, Imagine, CoralBay, Veset, etc.)
  • Experience building strong operational readiness practices (runbooks, alert tuning, on-call health, incident reviews)
  • Ability to create user interface designs based on client workflows

Additional Requirements:

Hybrid: This position currently has a hybrid schedule, which requires contributing from the office a minimum of four days per week. The Company reserves the right to change in-office requirements at any time.

As part of our selection process, external candidates may be required to attend an in-person interview with an NBCUniversal employee at one of our locations prior to a hiring decision. NBCUniversal's policy is to provide equal employment opportunities to all applicants and employees without regard to race, color, religion, creed, gender, gender identity or expression, age, national origin or ancestry, citizenship, disability, sexual orientation, marital status, pregnancy, veteran status, membership in the uniformed services, genetic information, or any other basis protected by applicable law.

If you are a qualified individual with a disability or a disabled veteran and require support throughout the application and/or recruitment process as a result of your disability, you have the right to request a reasonable accommodation. You can submit your request to AccessibilitySupport@nbcuni.com.

Job Location
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer, Playout
Sr. Site Reliability Engineer, Playout

NBCUniversal • Centennial (CO)

Hybrid
USD 130,000 - 160,000
Hybrid work schedule
Medical insurance
Dental insurance
+2
Sr. Site Reliability Engineer, Playout
Sr. Site Reliability Engineer, Playout

Worky • Centennial (CO)

Hybrid
USD 130,000 - 160,000
Medical insurance
Dental insurance
Vision insurance
+3
Staff Broadcast Reliability Engineer
Staff Broadcast Reliability Engineer

NBCUniversal • Colorado

Hybrid
USD 105,000 - 130,000
Hybrid schedule
Comprehensive benefits
Staff Broadcast Reliability Engineer
Staff Broadcast Reliability Engineer

NBCUniversal • Centennial (CO)

Hybrid
USD 105,000 - 130,000
Medical insurance
Dental insurance
Vision insurance
+3
Staff SRE: Live Playout & Cloud Reliability Leader
Staff SRE: Live Playout & Cloud Reliability Leader

NBCUniversal • Stamford (CT)

Hybrid
USD 145,000 - 175,000
Site Reliability Engineer
Site Reliability Engineer

NBCUniversal • Centennial (CO)

Hybrid
USD 110,000 - 145,000
Medical, dental, and vision insurance
401(k)
Paid leave
+1
Production Reliability Engineer
Production Reliability Engineer

NBCUniversal • New York (NY)

On-site
USD 110,000 - 130,000
Medical, dental, and vision insurance
401(k) plan
Paid leave
+2
Operations Supervisor
Operations Supervisor

NBCUniversal • Stamford (CT)

On-site
USD 70,000 - 95,000
Medical, dental, and vision insurance
401(k)
Tuition reimbursement
+2
Sr. Systems (Projects) Engineer - NBC4 &T44
Sr. Systems (Projects) Engineer - NBC4 &T44

NBCUniversal • Washington

On-site
USD 110,000 - 120,000
Medical, dental, vision insurance
401(k)
Tuition reimbursement
Master Control Operator
Master Control Operator

NBCUniversal • Stamford (CT)

On-site
USD 55,000 - 70,000
Medical, dental, and vision insurance
401(k)
Paid leave
+2