Job Title
Staff Software Engineer for Cluster Engineering
About the Role
We are hiring a Staff Software Engineer within the DTC Operational Engineering group with a focus on Cluster Engineering. This team builds a secure, efficient and easy‑to‑use container runtime platform and automates its scale across new markets, tenants and regions.
Responsibilities
As a staff software engineer on this team, you will be a subject‑matter expert in Kubernetes, helping to define the overall architecture and driving implementations. Your tasks include:
- Operating a large fleet of Kubernetes clusters (e.g., version upgrades, security patching, cost optimization, scaling and performance tuning).
- Identifying automation opportunities to reduce toil, especially during upgrades of hundreds of clusters ranging up to a thousand nodes.
- Leading cost optimization efforts such as tweaking Karpenter node‑scaling strategies, improving bin‑packing efficiency of pods, selecting appropriate node family types (m, c, r, spot and Graviton/ARM) and identifying over‑scaled infrastructure.
- Defining automation processes for creating new Kubernetes clusters, bootstrapping critical platform capabilities (service mesh, deployment systems, observability stack including metrics, tracing and logging).
- Working closely with SRE and other platform teams to improve reliability and efficiency of the platform.
- Managing lifecycle of critical platform components, setting quality standards and ensuring security and cost‑effectiveness.
Preferred Skills
- At least 8 years of software development, infrastructure management or operations experience.
- At least 5 years of Kubernetes and AWS experience.
- Experience with architecture design and developing platform solutions.
- Ability to mentor senior engineers.
- Experience with leading large technical projects with unclear requirements.
- Ability to drive decision making and cross‑team collaboration.
- Ability to write quality code for automation in Python, Bash or Go.
- Experience with designing infrastructure CI/CD pipelines (e.g., Jenkins or GitHub Actions).
- Experience with IaC, preferably Terraform.
- Used to Helm templating for Kubernetes manifests.
- Understanding of how GitOps tooling like ArgoCD or Flux works.
- Experience in rolling out infrastructure changes to production following a change‑management workflow.
- Knowing what metrics to monitor during a rollout to identify problems.
- Strong ownership mentality during rollouts, with capability to rollback and resolve issues.
- Strong sense of security, always using least privileges access and firewall configurations when needed.
- Understanding of how running workloads on Kubernetes clusters may be affected by cluster changes or node rotations.
- Willingness to talk to service development teams and understand their challenges during maintenance windows.
- Ability to define and measure KPIs and honor SLAs for infrastructure maintenance.
- Experience with Git and GitHub PR workflows.
- Experience working with Agile – Sprints, Epics/Stories, Jira.
- Experience building internal developer platforms using technologies such as Backstage/Argo Workflows.
Championing Inclusion at WBD
Warner Bros. Discovery embraces the opportunity to build a workforce that reflects a wide array of perspectives, backgrounds and experiences. Being an equal‑opportunity employer means that we consider qualified candidates on the basis of merit, regardless of sex, gender identity, ethnicity, age, sexual orientation, religion or belief, marital status, pregnancy, parenthood, disability or any other category protected by law.
If you’re a qualified candidate with a disability and you require adjustments or accommodations during the job application and/or recruitment process, please visit our accessibility page for instructions to submit your request.