Senior Cloud Reliability Engineer — Automation & Scale

Zilliz

Redwood City (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Zilliz is a fast-growing startup building a leading vector database for enterprise AI. We are expanding our Cloud Platform team to scale Zilliz Cloud with high ownership and automation.

You will own reliability end-to-end, debug complex issues across Kubernetes and cloud infra, and drive tooling that reduces toil while improving performance and scalability. Join a globally distributed team in a fast-paced environment, collaborating with database and infrastructure engineers to deliver reliable,

Qualifications

  • 3+ years building or operating production cloud systems, infrastructure platforms, database systems, or large-scale online services.
  • Bachelor's degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience.
  • Hands-on experience with Kubernetes, Docker, and at least one major cloud platform (AWS, GCP, or Azure).
  • Solid understanding of distributed systems; availability, scalability, performance, failure recovery, and operational tradeoffs.
  • Experience with distributed databases, storage systems, search systems, or large-scale online systems is a strong plus.
  • Experience operating highly multi-tenant systems or large infrastructure fleets; thousands of nodes, clusters, tenants, or customer deployments is especially valuable.
  • Familiarity with modern cloud operations tooling such as Terraform, Helm, Argo CD, Prometheus, Grafana, and CI/CD systems.
  • Strong bias for action, and the drive to thrive in a fast-paced, rapidly scaling environment.

Responsibilities

  • Own the reliability, availability, and production stability of Zilliz Cloud as we scale through the next stage of growth
  • Debug complex production issues across Kubernetes, cloud infrastructure, networking, storage, and distributed database systems
  • Build automation and diagnostic tooling; log analysis, alert correlation, incident investigation, runbook automation, and remediation workflows so problems get solved once, not repeatedly
  • Turn recurring incidents into reusable tools, automation, documentation, and product improvements
  • Improve observability across latency, availability, throughput, and resource efficiency
  • Partner with database and infrastructure engineers to make Zilliz Cloud more reliable, scalable, and automated

Skills

Kubernetes
Docker
Cloud platforms
Distributed systems
Observability

Education

Bachelor's degree in Computer Science/Software Engineering or related field

Tools

Terraform
Helm
Argo CD
Prometheus
Grafana
CI/CD

Job description

Zilliz is a fast-growing startup building a leading vector database for enterprise AI. We are expanding our Cloud Platform team to scale Zilliz Cloud with high ownership and automation.

You will own reliability end-to-end, debug complex issues across Kubernetes and cloud infra, and drive tooling that reduces toil while improving performance and scalability. Join a globally distributed team in a fast-paced environment, collaborating with database and infrastructure engineers to deliver reliable,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Cloud Reliability
Senior Software Engineer, Cloud Reliability

Zilliz • Redwood City (CA)

On-site
USD 140,000 - 210,000
Senior Software Engineer, Cloud Platform
Senior Software Engineer, Cloud Platform

Zilliz • Redwood City (CA)

On-site
USD 175,000 - 225,000
Competitive compensation (cash + equity)
Medical, dental, and vision insurance
Generous 401(k) and regional retirement plans
Senior Site Reliability Engineer, Cloud Platform
Senior Site Reliability Engineer, Cloud Platform

Zilliz • Redwood City (CA)

On-site
USD 175,000 - 225,000
Senior Cloud Platform Engineer - AI-First, Multi-Cloud
Senior Cloud Platform Engineer - AI-First, Multi-Cloud

Medium • Redwood City (CA)

Hybrid
USD 175,000 - 225,000
Competitive compensation (cash + equity)
Medical, dental, and vision insurance
Generous 401(k) and regional retirement plans
Cloud Reliability Engineer | Scale, Predict, Deliver
Cloud Reliability Engineer | Scale, Predict, Deliver

ZT Group Intl, Inc. dba ZT Systems • Secaucus (NJ)

On-site
USD 105,000 - 154,000
Staff Platform Infra Engineer — Scale & Reliability
Staff Platform Infra Engineer — Scale & Reliability

Zūm • Redwood City (CA)

On-site
USD 240,000 - 280,000
Medical
Dental
Vision
+5
Platform Reliability Engineer – Cloud & Automation
Platform Reliability Engineer – Cloud & Automation

startupjobs.pt • United States

On-site
USD 130,000 - 190,000
Senior Infrastructure Lead | AI-Driven Cloud & Automation
Senior Infrastructure Lead | AI-Driven Cloud & Automation

Zelis • Atlanta (GA)

Hybrid
USD 102,000 - 140,000
401(k) with employer match
Paid time off
Health insurance
+2
Cloud Database Reliability Engineer: Automation & IaC
Cloud Database Reliability Engineer: Automation & IaC

Alexander Technology Group • Dedham (MA)

On-site
USD 140,000 - 190,000
Cloud Infrastructure Engineer: Scale, Reliability & Automation
Cloud Infrastructure Engineer: Scale, Reliability & Automation

People Culture Talent • Chicago (IL)

On-site
USD 110,000 - 160,000