Senior Datacenter Automation Engineer (Kubernetes & Infra)

Zipline

South San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Zipline is seeking a Senior Software Engineer for the Infrastructure team in South San Francisco. You will own the datacenter compute and storage lifecycle, from bare-metal provisioning to Kubernetes lifecycle, while shaping automation to speed deployments and reduce manual work.

You will drive reliability with SLIs/SLOs, lead incident management for flight-critical infra, and implement monitoring, capacity planning, and cost controls across the global platform.

Qualifications

  • 5+ years of engineering experience with at least 4 years owning production datacenter, virtualization, or infrastructure automation systems.
  • Deep, hands-on expertise with bare-metal provisioning and imaging (PXE/iPXE, IPMI, Redfish), hypervisors (KVM/qemu, ESXi or equivalent), and storage systems (Ceph, NVMeoF, SAN) at scale.
  • Proven Kubernetes operations experience: cluster provisioning, upgrades, control-plane HA, kubeadm/cluster API or equivalent, CNI and CSI troubleshooting, and workload scheduling at multi-cluster scale.
  • Production-grade automation and coding skills in one or more languages (Python, Go, or Rust) and experience with CI/CD pipelines, Terraform/Ansible/Helm, and GitOps practices.
  • Strong networking fundamentals: VLANs, BGP/EVPN at leaf/spine, LACP, routing, and network troubleshooting for cluster networking and storage fabrics.
  • On-call and incident experience: you have owned postmortems, SLIs/SLOs, and driven reliability improvements under operational pressure.
  • Physical datacenter readiness: able to work on-site in South San Francisco HQ with regular in-office cadence, plus occasional travel to partner datacenters or field sites and hands-on rack/cable/repair work when required.
  • Security and safety mindset: experience operating in regulated or safety-sensitive environments, following change control and audit processes.
  • Clear communication and cross-team ownership: you will partner with flight software, field ops, hardware, and SRE teams and must translate operational needs into automated, testable systems.

Responsibilities

  • Own end-to-end lifecycle for datacenter compute and storage: bare-metal provisioning, hypervisor management, SAN/NVMe storage clusters, network configuration, and Kubernetes cluster lifecycle.
  • Design, build, and operate automation that reduces manual setup time and increases deployment velocity: PXE/firmware workflows, dynamic inventory, image generation, fleet-wide configuration drift detection, and automated recovery playbooks.
  • Deliver measurable reliability and scale improvements: set SLIs/SLOs for provisioning time, node commissioning success rate, cluster upgrade success rate, and mean time to recover (MTTR); own meeting those targets.
  • Lead cross-functional runbook and incident ownership for infra incidents affecting flight operations or telemetry: on-call rotation, incident commander for datacenter platform incidents, postmortems and action items.
  • Instrument and maintain monitoring, alerting, and dashboards for hardware health, hypervisor performance, storage latency, Kubernetes control plane health, and cluster autoscaling behavior.
  • Implement cost, capacity, and lifecycle management: capacity planning for compute/storage, automated reclamation, firmware/BIOS/hypervisor patch pipelines, and cold-standby / failover procedures for critical systems.
  • Execute hands-on tasks when required: racking and cabling in datacenters, troubleshooting hardware failures, capture forensic logs, and coordinate physical repairs with vendors and field ops.

Skills

Datacenter infrastructure
Bare-metal provisioning
KVM/qemu virtualization
ESXi
Ceph
NVMeoF
Kubernetes operations
CI/CD pipelines
Python
Go
Rust
Terraform
Ansible
Helm
GitOps
Networking fundamentals
VLANs/BGP/EVPN
On-call/incidents

Tools

PXE/iPXE
IPMI/Redfish
KVM
SAN
Hypervisor management

Job description

Zipline is seeking a Senior Software Engineer for the Infrastructure team in South San Francisco. You will own the datacenter compute and storage lifecycle, from bare-metal provisioning to Kubernetes lifecycle, while shaping automation to speed deployments and reduce manual work.

You will drive reliability with SLIs/SLOs, lead incident management for flight-critical infra, and implement monitoring, capacity planning, and cost controls across the global platform.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer, Cloud Data Platform
Senior Data Engineer, Cloud Data Platform

Namely • South San Francisco (CA)

On-site
USD 140,000 - 190,000
Senior IT Infra & IAM Engineer — Automation & Cloud
Senior IT Infra & IAM Engineer — Automation & Cloud

Namely • South San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior IT Infra and Identity Management Engineer
Senior IT Infra and Identity Management Engineer

Zipline • South San Francisco (CA)

On-site
USD 150,000 - 190,000
Sr. SWE Datacenter Automation
Sr. SWE Datacenter Automation

Zipline • South San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior DevOps Engineer - AI Infra, Kubernetes & Cloud
Senior DevOps Engineer - AI Infra, Kubernetes & Cloud

StackAI • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Hybrid work model
Remote work available
Office near Salesforce Park
Senior Analytics Platform Engineer
Senior Analytics Platform Engineer

Zipline • San Francisco (CA)

On-site
USD 190,000 - 230,000
Senior Cloud & Kubernetes Infrastructure Engineer
Senior Cloud & Kubernetes Infrastructure Engineer

Skydio, Inc. • San Mateo (CA)

On-site
USD 190,000 - 250,000
Equity (stock options)
Health insurance
401K savings plan
+1
Senior Software Engineer, Airspace Platform
Senior Software Engineer, Airspace Platform

Zipline International Inc. • South San Francisco (CA)

On-site
USD 140,000 - 200,000
Strategic Operations Analyst, Data & Automation
Strategic Operations Analyst, Data & Automation

Zipline • South San Francisco (CA)

On-site
USD 130,000 - 170,000
Equity compensation
Medical, dental and vision insurance
Paid time off
Senior Analytics Engineer: Production Data Platform Leader
Senior Analytics Engineer: Production Data Platform Leader

Zipline • South San Francisco (CA)

On-site
USD 170,000 - 230,000