CoreWeave logo

Staff Software Engineer (Cluster Orchestration)

CoreWeave

On-siteSunnyvale, CAlead$185k–$275kPosted 9h ago

Job description

  • As part of the Cluster Orchestration team, you will play a key role in advancing CoreWeave’s orchestration platform including SUNK (Slurm on Kubernetes) and beyond, our Kubernetes-native foundation that powers AI training and inference at scale

  • This is an opportunity to help shape one of the most critical layers of the AI cloud: ensuring workloads run seamlessly, reliably, and efficiently across massive GPU clusters

  • By building the systems that eliminate infrastructure bottlenecks and create new orchestration capabilities, you will directly empower customers to innovate faster and push the boundaries of what’s possible with AI

  • As a Staff Engineer, you will be a technical leader shaping the long-term strategy for CoreWeave’s orchestration platform

  • You’ll define architectural direction, own critical parts of the orchestration platform and other managed services, and drive cross-org initiatives in scheduling, quota enforcement, and scaling at hyperscale

  • You’ll mentor senior engineers, establish org-wide best practices in reliability and observability, and ensure CoreWeave’s orchestration layer evolves to meet the demands of next-generation AI workloads- 8+ years of software engineering experience

  • Proven track record designing and operating large-scale distributed systems in production

  • Experience setting technical direction and influencing cross-team architecture

  • Advanced proficiency in Go and distributed systems design and cloud-native development

  • Deep expertise in Slurm/Kubernetes internals and cloud-native development

  • Bachelor’s or Master’s degree in CS, EE, or related field

  • Deep expertise in Slurm/Kubernetes internals

  • Familiarity with orchestration and workflow technologies such as Ray, Kubeflow, Kueue, Istio, Knative, or Argo Workflows

  • Experience with distributed workloads, GPU-based applications, or ML pipelines

  • Knowledge of scheduling concepts like quota enforcement, pre-emption, and scaling strategies

  • Exposure to reliability practices including SLOs, alarms, and post-incident reviews

  • Experience with AI infrastructure and workloads (ML training, inference, or HPC)

  • Ability to mentor senior engineers and elevate organizational standards

  • You thrive on solving problems that balance cost, performance, and reliability in high-demand environments

  • You’re curious about orchestration beyond SUNK and how to evolve it for next-generation AI

  • You love defining long-term architecture for systems at global scale

  • You’re an expert at mentorship, architecture, and operational excellence across teams