Zilliz logo

Senior Software Engineer (Cloud Reliability)

Zilliz

Remotesenior$175k–$225kPosted 8h ago

Job description

  • We’re entering our next phase of 10x growth; more customers, larger datasets, and far higher expectations for reliability

  • You’ll join a small, fast-moving Cloud Platform team that operates large-scale, multi-cloud, distributed database systems in production

  • This is a high-ownership role for engineers who want to move fast, build automation instead of toil, and take real responsibility for production stability

  • Own the reliability, availability, and production stability of Zilliz Cloud as we scale through the next stage of growth

  • Debug complex production issues across Kubernetes, cloud infrastructure, networking, storage, and distributed database systems

  • Build automation and diagnostic tooling; log analysis, alert correlation, incident investigation, runbook automation, and remediation workflows so problems get solved once, not repeatedly

  • Turn recurring incidents into reusable tools, automation, documentation, and product improvements

  • Improve observability across latency, availability, throughput, and resource efficiency

  • Partner with database and infrastructure engineers to make Zilliz Cloud more reliable, scalable, and automated

Benefits

  • Equity

  • Regular bonus and equity refresh opportunities

  • Comprehensive medical, dental, and vision insurance

  • Paid time off, including vacation, bereavement, and sick days

  • Generous 401(k) and regional retirement plans

  • Hybrid work model/Remote work opportunities available- Bachelor’s degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience

  • 3+ years building or operating production cloud systems, infrastructure platforms, database systems, or large-scale online services

  • Experience operating highly multi-tenant systems or large infrastructure fleets; thousands of nodes, clusters, tenants, or customer deployments is especially valuable

  • Strong bias for action, and the drive to thrive in a fast-paced, rapidly scaling environment

  • Experience with distributed databases, storage systems, search systems, or large-scale online systems is a strong plus

  • Solid understanding of distributed systems; availability, scalability, performance, failure recovery, and operational tradeoffs

  • Familiarity with modern cloud operations tooling such as Terraform, Helm, Argo CD, Prometheus, Grafana, and CI/CD systems

  • Strong hands-on experience with Kubernetes, Docker, and at least one major cloud platform (AWS, GCP, or Azure)