
Senior Software Engineer (Cloud Reliability)
Zilliz
Job description
-
We’re entering our next phase of 10x growth; more customers, larger datasets, and far higher expectations for reliability
-
You’ll join a small, fast-moving Cloud Platform team that operates large-scale, multi-cloud, distributed database systems in production
-
This is a high-ownership role for engineers who want to move fast, build automation instead of toil, and take real responsibility for production stability
-
Own the reliability, availability, and production stability of Zilliz Cloud as we scale through the next stage of growth
-
Debug complex production issues across Kubernetes, cloud infrastructure, networking, storage, and distributed database systems
-
Build automation and diagnostic tooling; log analysis, alert correlation, incident investigation, runbook automation, and remediation workflows so problems get solved once, not repeatedly
-
Turn recurring incidents into reusable tools, automation, documentation, and product improvements
-
Improve observability across latency, availability, throughput, and resource efficiency
-
Partner with database and infrastructure engineers to make Zilliz Cloud more reliable, scalable, and automated
Benefits
-
Equity
-
Regular bonus and equity refresh opportunities
-
Comprehensive medical, dental, and vision insurance
-
Paid time off, including vacation, bereavement, and sick days
-
Generous 401(k) and regional retirement plans
-
Hybrid work model/Remote work opportunities available- Bachelor’s degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience
-
3+ years building or operating production cloud systems, infrastructure platforms, database systems, or large-scale online services
-
Experience operating highly multi-tenant systems or large infrastructure fleets; thousands of nodes, clusters, tenants, or customer deployments is especially valuable
-
Strong bias for action, and the drive to thrive in a fast-paced, rapidly scaling environment
-
Experience with distributed databases, storage systems, search systems, or large-scale online systems is a strong plus
-
Solid understanding of distributed systems; availability, scalability, performance, failure recovery, and operational tradeoffs
-
Familiarity with modern cloud operations tooling such as Terraform, Helm, Argo CD, Prometheus, Grafana, and CI/CD systems
-
Strong hands-on experience with Kubernetes, Docker, and at least one major cloud platform (AWS, GCP, or Azure)