Senior Staff Engineer - Kubernetes Platform and Networking

Equinix

Bengaluru

On-site

INR 4,000,000 - 6,800,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Equinix is seeking a Senior Staff Engineer to lead the Kubernetes platform and networking reliability across multiple metros. You will own OS provisioning, cluster bootstrapping, and the automation that enables zero-touch deployment.

You will partner with hardware, underlay network, and data plane teams, shaping design in practice while maintaining rigorous runbooks and blameless postmortems.

Qualifications

  • Deep hands-on Linux/Ubuntu provisioning at scale with PXE/iPXE imaging.
  • Kubernetes deployment and operation on bare metal, HA and multi-cluster fleet management.
  • CI/CD for infrastructure using IaC and GitOps.
  • Experience with networking, underlay, VLANs, BGP and private connectivity.

Responsibilities

  • Provision OS across bare-metal servers in multiple metros.
  • Deploy and operate Kubernetes clusters with HA and correct fault domain placement.
  • Own CI/CD pipeline that deploys the data plane and enforce validation gates.
  • Manage etcd snapshot/restore procedures and upgrade runbooks.
  • Drive GitOps-based fleet configuration management and incident response.

Skills

Linux provisioning
Kubernetes
RKE2 / kubeadm
CI/CD for infra
Automation

Tools

Rancher Prime
ArgoCD
Terraform
Ansible
Fleet

Job description

Senior Staff Engineer, Kubernetes Platform & Networking

Who are we? Equinix is the world’s digital infrastructure company®, shortening the path to connectivity to enable the innovations that enrich our work, life and planet. A place where bold ideas are welcomed, human connection is valued, and everyone has the opportunity to shape their future. Help us challenge assumptions, uncover bias, and remove barriers—because progress starts with fresh ideas. You’ll find belonging, purpose, and a team that welcomes you—because when you feel valued, you’re empowered to do your best work.

Job Summary

We are looking for a Senior Staff Engineer to join the Reliability Engineering team that operates a bare‑metal Kubernetes platform across multiple metro locations. The platform’s architecture is defined and the build is underway, you will help deploy, operate and continuously improve it, from OS provisioning on racked servers through a multi‑cluster fleet managed from a central management plane, with a software‑defined data plane deployed via GitOps. This is a deliberately broad role. We are not looking for someone who owns a layer and hands everything below it across a boundary. You will work alongside the teams that own the physical hardware, the underlay network, and the data plane software, and you need enough fluency in all three to reason about a fault before escalating it. Where you find the design falls short in practice, we expect you to say so and to shape how it evolves. You will join the Digital Interconnection Engineering organization. The team owns the full application stack for interconnection and the Kubernetes platform for the software‑defined data plane across multiple metros. We operate with an automation‑first, GitOps‑driven culture and a strong commitment to operational rigor, documented runbooks, and blameless postmortems.

Responsibilities
OS Provisioning & Bare‑Metal Infrastructure

Own OS installation (Debian/Ubuntu) across bare‑metal servers in multiple metros using PXE/iPXE imaging and cloud‑init, and harden the OS prior to cluster bootstrap

Build and maintain a repeatable, zero‑touch provisioning pipeline so every metro is deployed consistently without manual per‑node intervention

Partner with the teams that own the hardware on BIOS/firmware baselines, high‑speed NIC configuration and disk layout standards (separate NVMe for etcd, SSD for OS) and be able to diagnose faults at that layer, not just report them

Kubernetes Control Plane & Worker Nodes

Deploy and operate Kubernetes clusters across all metros with quorum‑based control‑plane fault tolerance and correct failure‑domain placement across racks, PDUs, or availability zones

Maintain the Layer 4 load balancer or keepalived VIP fronting all kube‑apiserver instances, so no client depends on a single node

Manage worker node provisioning, CNI configuration, and end‑to‑end cluster validation before workloads are onboarded

CI/CD & Data Plane Deployment

Build and own the CI/CD pipeline that deploys the software‑defined data plane, including automated throughput and latency validation gates that must pass before promotion to production. The data plane software itself is configured and troubleshot by the team that owns it, your ownership is the pipeline and the gates

Own data‑plane observability: monitoring, alerting thresholds, and operational runbooks

Network

Operate on top of a multi‑metro underlay and partner with network and security teams on VLANs, BGP peering, NIC bonding, private connectivity, firewall rules and segmentation

Validate connectivity end to end across environments, and troubleshoot far enough into the network layer to isolate a problem before handing it over

OS & Kubernetes Upgrades

Plan and execute rolling OS and Kubernetes upgrades across the fleet with zero unplanned downtime; maintain tested upgrade runbooks and validated etcd snapshot/restore procedures before every major change

Own the etcd snapshot schedule, retention policy, and offsite storage; verify restore runbooks before each site goes live

Apply the platform’s patching cadence across the fleet, tracking CVEs affecting the OS, Kubernetes and cluster components

Management Plane

Administer the centralized management plane across all clusters, including RBAC, cluster registration and policy enforcement, and keep production and non‑production management planes strictly isolated to prevent blast‑radius cascades from configuration changes

Drive GitOps‑based fleet configuration management to keep clusters consistent, auditable, and drift‑free

On‑Call & Incident Management

Own your shift of a follow‑the‑sun on‑call rotation, and lead incident response for control‑plane, data‑plane, and network events

Drive blameless postmortems and systematically eliminate recurring failure modes through automation, better runbooks, and sharper alerting

Contribute to the shared incident‑command process, and maintain the runbooks for the systems you operate

What Success Looks Like, First 9-12 Months
  • Provisioning is fully automated and repeatable, new nodes and metros brought up with no manual per‑node steps
  • Deployment pipeline in production, with throughput and latency gates enforced before promotion
  • Management plane administered across all clusters; production and non‑production planes isolated; GitOps driving cluster configuration
  • OS and Kubernetes upgrades executed on cadence across the fleet with documented, tested runbooks
  • etcd snapshot and restore tested and signed off before each site goes live
  • Observability and alerting in place for the systems you operate, including the data plane
  • Recurring failure modes measurably reduced through automation and runbook improvements
Qualifications
  • OS & bare metal: Deep hands‑on Linux/Ubuntu provisioning at scale, PXE/iPXE imaging, cloud‑init, BIOS/firmware management, OS hardening
  • Hardware fluency: Comfortable with BMC/IPMI, SMART data and platform sensors; able to diagnose disk, memory, NIC and power faults and work with vendors through warranty replacement
  • Kubernetes: Strong experience deploying and operating RKE2 or kubeadm clusters on bare metal, etcd operations, kube‑apiserver HA, CNI (Multus/SR‑IOV), multi‑cluster fleet management
  • Rancher or equivalent management plane: Production experience managing multiple clusters through Rancher Prime, Rancher, or comparable fleet tooling
  • CI/CD for infrastructure: GitOps or pipeline‑driven infrastructure deployment, Ansible, Terraform, ArgoCD, or Fleet
  • Networking: Hands‑on with VLANs, BGP, bonded NICs, SR‑IOV and private connectivity, with enough depth to troubleshoot across the boundary rather than only consume it
  • Upgrades: Demonstrated zero‑downtime rolling upgrades of both OS and Kubernetes across a multi‑node fleet
  • Operational rigor: Experience running platforms where availability is measured in minutes of downtime per year, and where an unreviewed change is the most likely cause of an outage
  • On‑call: Comfortable owning an on‑call shift and leading infrastructure incident response
  • Automation mindset: Everything‑as‑code, repeatable and auditable; strong bias toward automation over manual operations
  • Preferred: Familiarity with 6WIND VSR or other DPDK‑based data planes, enough to build and validate a deployment and performance‑testing pipeline around one SR‑IOV NIC tuning and data‑plane performance validation
  • SDN or private‑connectivity platforms (Equinix Fabric, VyOS, or equivalent)
  • Familiarity with LLM provider APIs and AI/agent gateway proxying patterns, since inference traffic from multiple providers transits this platform
  • Experience operating infrastructure across geographically distributed sites

Equinix is committed to ensuring that our employment process is open to all individuals, including those with a disability.

Equinix is an Equal Employment Opportunity and, in the U.S., an Aff

All qualified applicants will receive consideration for employment without regard to unlawful consideration of race, color, religion, creed, national or ethnic origin, ancestry, place of birth, citizenship, sex, pregnancy / childbirth or related medical conditions, sexual orientation, gender identity or expression, marital or domestic partnership status, age, veteran or military status, physical or mental disability, medical condition, genetic information, political / organizational affiliation, status as a victim or family member of a victim of crime or abuse, or any other status protected by applicable law.

We use artificial intelligence in our hiring process.

This posting is a new position within our organization.

Experience Level Senior Level

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Kubernetes Platform and Networking Engineer
Senior Kubernetes Platform and Networking Engineer

Equinix • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Senior Staff Engineer, Kubernetes Platform & Networking
Senior Staff Engineer, Kubernetes Platform & Networking

Equinix • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Staff Engineer, Kubernetes Platform & Networking
Staff Engineer, Kubernetes Platform & Networking

Equinix, Inc. • Bengaluru

On-site
INR 350,000 - 550,000
Staff Engineer, Kubernetes Platform & Networking
Staff Engineer, Kubernetes Platform & Networking

Equinix • Bengaluru

On-site
INR 400,000 - 700,000
Staff Software Engineer - Cloud Platform & DevSecOps
Staff Software Engineer - Cloud Platform & DevSecOps

Equinix • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Staff Software Engineer - Cloud Platform & DevSecOps
Staff Software Engineer - Cloud Platform & DevSecOps

Equinix, Inc. • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Staff Cloud Platform Engineer
Senior Staff Cloud Platform Engineer

Equinix • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Engineering Manager - Cloud Platform and DevSecOps
Engineering Manager - Cloud Platform and DevSecOps

Equinix • Bengaluru

On-site
INR 4,000,000 - 8,000,000
Engineering Manager – Cloud Platform & DevSecOps
Engineering Manager – Cloud Platform & DevSecOps

Equinix • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Engineering Manager – Cloud Platform & DevSecOps
Engineering Manager – Cloud Platform & DevSecOps

Equinix, Inc. • Bengaluru

On-site
INR 6,000,000 - 9,000,000