Get more replies from employers
Send a job-specific resume in minutes.
Equinix is seeking a Senior Staff Engineer to lead the Kubernetes platform and networking reliability across multiple metros. You will own OS provisioning, cluster bootstrapping, and the automation that enables zero-touch deployment.
You will partner with hardware, underlay network, and data plane teams, shaping design in practice while maintaining rigorous runbooks and blameless postmortems.
Who are we? Equinix is the world’s digital infrastructure company®, shortening the path to connectivity to enable the innovations that enrich our work, life and planet. A place where bold ideas are welcomed, human connection is valued, and everyone has the opportunity to shape their future. Help us challenge assumptions, uncover bias, and remove barriers—because progress starts with fresh ideas. You’ll find belonging, purpose, and a team that welcomes you—because when you feel valued, you’re empowered to do your best work.
We are looking for a Senior Staff Engineer to join the Reliability Engineering team that operates a bare‑metal Kubernetes platform across multiple metro locations. The platform’s architecture is defined and the build is underway, you will help deploy, operate and continuously improve it, from OS provisioning on racked servers through a multi‑cluster fleet managed from a central management plane, with a software‑defined data plane deployed via GitOps. This is a deliberately broad role. We are not looking for someone who owns a layer and hands everything below it across a boundary. You will work alongside the teams that own the physical hardware, the underlay network, and the data plane software, and you need enough fluency in all three to reason about a fault before escalating it. Where you find the design falls short in practice, we expect you to say so and to shape how it evolves. You will join the Digital Interconnection Engineering organization. The team owns the full application stack for interconnection and the Kubernetes platform for the software‑defined data plane across multiple metros. We operate with an automation‑first, GitOps‑driven culture and a strong commitment to operational rigor, documented runbooks, and blameless postmortems.
Own OS installation (Debian/Ubuntu) across bare‑metal servers in multiple metros using PXE/iPXE imaging and cloud‑init, and harden the OS prior to cluster bootstrap
Build and maintain a repeatable, zero‑touch provisioning pipeline so every metro is deployed consistently without manual per‑node intervention
Partner with the teams that own the hardware on BIOS/firmware baselines, high‑speed NIC configuration and disk layout standards (separate NVMe for etcd, SSD for OS) and be able to diagnose faults at that layer, not just report them
Deploy and operate Kubernetes clusters across all metros with quorum‑based control‑plane fault tolerance and correct failure‑domain placement across racks, PDUs, or availability zones
Maintain the Layer 4 load balancer or keepalived VIP fronting all kube‑apiserver instances, so no client depends on a single node
Manage worker node provisioning, CNI configuration, and end‑to‑end cluster validation before workloads are onboarded
Build and own the CI/CD pipeline that deploys the software‑defined data plane, including automated throughput and latency validation gates that must pass before promotion to production. The data plane software itself is configured and troubleshot by the team that owns it, your ownership is the pipeline and the gates
Own data‑plane observability: monitoring, alerting thresholds, and operational runbooks
Operate on top of a multi‑metro underlay and partner with network and security teams on VLANs, BGP peering, NIC bonding, private connectivity, firewall rules and segmentation
Validate connectivity end to end across environments, and troubleshoot far enough into the network layer to isolate a problem before handing it over
Plan and execute rolling OS and Kubernetes upgrades across the fleet with zero unplanned downtime; maintain tested upgrade runbooks and validated etcd snapshot/restore procedures before every major change
Own the etcd snapshot schedule, retention policy, and offsite storage; verify restore runbooks before each site goes live
Apply the platform’s patching cadence across the fleet, tracking CVEs affecting the OS, Kubernetes and cluster components
Administer the centralized management plane across all clusters, including RBAC, cluster registration and policy enforcement, and keep production and non‑production management planes strictly isolated to prevent blast‑radius cascades from configuration changes
Drive GitOps‑based fleet configuration management to keep clusters consistent, auditable, and drift‑free
Own your shift of a follow‑the‑sun on‑call rotation, and lead incident response for control‑plane, data‑plane, and network events
Drive blameless postmortems and systematically eliminate recurring failure modes through automation, better runbooks, and sharper alerting
Contribute to the shared incident‑command process, and maintain the runbooks for the systems you operate
Equinix is committed to ensuring that our employment process is open to all individuals, including those with a disability.
Equinix is an Equal Employment Opportunity and, in the U.S., an Aff
All qualified applicants will receive consideration for employment without regard to unlawful consideration of race, color, religion, creed, national or ethnic origin, ancestry, place of birth, citizenship, sex, pregnancy / childbirth or related medical conditions, sexual orientation, gender identity or expression, marital or domestic partnership status, age, veteran or military status, physical or mental disability, medical condition, genetic information, political / organizational affiliation, status as a victim or family member of a victim of crime or abuse, or any other status protected by applicable law.
We use artificial intelligence in our hiring process.
This posting is a new position within our organization.
Experience Level Senior Level