Senior Software Engineer, Cloud Reliability

Zilliz

Redwood City (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Zilliz is a fast-growing startup building a leading vector database for enterprise AI. We are expanding our Cloud Platform team to scale Zilliz Cloud with high ownership and automation.

You will own reliability end-to-end, debug complex issues across Kubernetes and cloud infra, and drive tooling that reduces toil while improving performance and scalability. Join a globally distributed team in a fast-paced environment, collaborating with database and infrastructure engineers to deliver reliable,

Qualifications

  • 3+ years building or operating production cloud systems, infrastructure platforms, database systems, or large-scale online services.
  • Bachelor's degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience.
  • Hands-on experience with Kubernetes, Docker, and at least one major cloud platform (AWS, GCP, or Azure).
  • Solid understanding of distributed systems; availability, scalability, performance, failure recovery, and operational tradeoffs.
  • Experience with distributed databases, storage systems, search systems, or large-scale online systems is a strong plus.
  • Experience operating highly multi-tenant systems or large infrastructure fleets; thousands of nodes, clusters, tenants, or customer deployments is especially valuable.
  • Familiarity with modern cloud operations tooling such as Terraform, Helm, Argo CD, Prometheus, Grafana, and CI/CD systems.
  • Strong bias for action, and the drive to thrive in a fast-paced, rapidly scaling environment.

Responsibilities

  • Own the reliability, availability, and production stability of Zilliz Cloud as we scale through the next stage of growth
  • Debug complex production issues across Kubernetes, cloud infrastructure, networking, storage, and distributed database systems
  • Build automation and diagnostic tooling; log analysis, alert correlation, incident investigation, runbook automation, and remediation workflows so problems get solved once, not repeatedly
  • Turn recurring incidents into reusable tools, automation, documentation, and product improvements
  • Improve observability across latency, availability, throughput, and resource efficiency
  • Partner with database and infrastructure engineers to make Zilliz Cloud more reliable, scalable, and automated

Skills

Kubernetes
Docker
Cloud platforms
Distributed systems
Observability

Education

Bachelor's degree in Computer Science/Software Engineering or related field

Tools

Terraform
Helm
Argo CD
Prometheus
Grafana
CI/CD

Job description

Zilliz is a fast-growing startup developing the industry’s leading vector database for enterprise-grade AI. Founded by the engineers behind Milvus, the world’s most popular open-source vector database, the company builds next‑generation database technologies to help organizations quickly create AI applications. On a mission to democratize AI, Zilliz is committed to simplifying data management for AI applications and making vector databases accessible to every organization.

We’re entering our next phase of 10x growth; more customers, larger datasets, and far higher expectations for reliability. You’ll join a small, fast‑moving Cloud Platform team that operates large‑scale, multi‑cloud, distributed database systems in production. This is a high‑ownership role for engineers who want to move fast, build automation instead of toil, and take real responsibility for production stability.

What you will do:
  • Own the reliability, availability, and production stability of Zilliz Cloud as we scale through the next stage of growth
  • Debug complex production issues across Kubernetes, cloud infrastructure, networking, storage, and distributed database systems
  • Build automation and diagnostic tooling; log analysis, alert correlation, incident investigation, runbook automation, and remediation workflows so problems get solved once, not repeatedly
  • Turn recurring incidents into reusable tools, automation, documentation, and product improvements
  • Improve observability across latency, availability, throughput, and resource efficiency
  • Partner with database and infrastructure engineers to make Zilliz Cloud more reliable, scalable, and automated
What we are looking for:
  • 3+ years building or operating production cloud systems, infrastructure platforms, database systems, or large‑scale online services
  • Bachelor's degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience
  • Strong hands‑on experience with Kubernetes, Docker, and at least one major cloud platform (AWS, GCP, or Azure)
  • Solid understanding of distributed systems; availability, scalability, performance, failure recovery, and operational tradeoffs
  • Experience with distributed databases, storage systems, search systems, or large‑scale online systems is a strong plus
  • Experience operating highly multi‑tenant systems or large infrastructure fleets; thousands of nodes, clusters, tenants, or customer deployments is especially valuable
  • Familiarity with modern cloud operations tooling such as Terraform, Helm, Argo CD, Prometheus, Grafana, and CI/CD systems
  • Strong bias for action, and the drive to thrive in a fast‑paced, rapidly scaling environment
How we operate:
  • High ownership: You own production reliability end‑to‑end. The whole system, not a slice of it. High autonomy, high trust, minimal process
  • Fast and focused:We ship often and keep a high bar. This team suits engineers who want velocity and a steep growth curve over red tape
  • Globally distributed: We work closely with our core engineering teams across APAC. Occasional early morning or evening syncs in exchange for an on‑call setup designed around timezone coverage, not overnight pages

Zilliz is an Equal Opportunity Employer. We consider all qualified applicants for employment without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, genetic information, veteran status, or any other characteristic protected by applicable law.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Cloud Platform
Senior Software Engineer, Cloud Platform

Zilliz • Redwood City (CA)

On-site
USD 175,000 - 225,000
Competitive compensation (cash + equity)
Medical, dental, and vision insurance
Generous 401(k) and regional retirement plans
Senior Site Reliability Engineer, Cloud Platform
Senior Site Reliability Engineer, Cloud Platform

Zilliz • Redwood City (CA)

On-site
USD 175,000 - 225,000
Senior Cloud Reliability Engineer — Automation & Scale
Senior Cloud Reliability Engineer — Automation & Scale

Zilliz • Redwood City (CA)

On-site
USD 140,000 - 210,000
Lead Recruiter, US & Europe
Lead Recruiter, US & Europe

Zilliz • Redwood City (CA)

On-site
USD 150,000 - 190,000
Enterprise Account Executive - SF Bay Area
Enterprise Account Executive - SF Bay Area

Zilliz • Redwood City (CA)

On-site
USD 140,000 - 210,000
Competitive compensation
Regular bonus and equity refresh
Medical, dental, and vision insurance
+2
Senior Software Engineer, Vector Index Research
Senior Software Engineer, Vector Index Research

Zilliz • Redwood City (CA)

On-site
USD 175,000 - 250,000
Competitive compensation (cash + equity)
Medical, dental, and vision insurance
Paid time off
+1
Cloud Infrastructure Engineer
Cloud Infrastructure Engineer

Mixpeek • San Mateo (CA)

On-site
USD 180,000 - 280,000
Equity
Full benefits (Medical, Dental, Vision
401k
+2
Enterprise Account Executive - New York
Enterprise Account Executive - New York

Zilliz • New York (NY)

On-site
USD 100,000 - 175,000
Competitive compensation (cash + equity)
Regular bonus and equity refresh opportunities
Medical, dental, and vision insurance
+2
Enterprise Account Executive - Texas
Enterprise Account Executive - Texas

Zilliz • Austin (TX)

On-site
USD 100,000 - 175,000
Competitive compensation (cash + equity)
Regular bonus and equity refresh opportunities
Medical, dental, and vision insurance
+2
Enterprise Account Executive - Texas
Enterprise Account Executive - Texas

Zilliz • Austin (MO)

On-site
USD 100,000 - 175,000
Competitive compensation (cash + equity)
Regular bonus and equity refresh opportunities
Medical, dental, and vision insurance
+2