
Staff Software Engineer (Cluster Orchestration)
CoreWeave
Job description
-
As part of the Cluster Orchestration team, you will play a key role in advancing CoreWeave’s orchestration platform including SUNK (Slurm on Kubernetes) and beyond, our Kubernetes-native foundation that powers AI training and inference at scale
-
This is an opportunity to help shape one of the most critical layers of the AI cloud: ensuring workloads run seamlessly, reliably, and efficiently across massive GPU clusters
-
By building the systems that eliminate infrastructure bottlenecks and create new orchestration capabilities, you will directly empower customers to innovate faster and push the boundaries of what’s possible with AI
-
As a Staff Engineer, you will be a technical leader shaping the long-term strategy for CoreWeave’s orchestration platform
-
You’ll define architectural direction, own critical parts of the orchestration platform and other managed services, and drive cross-org initiatives in scheduling, quota enforcement, and scaling at hyperscale
-
You’ll mentor senior engineers, establish org-wide best practices in reliability and observability, and ensure CoreWeave’s orchestration layer evolves to meet the demands of next-generation AI workloads- 8+ years of software engineering experience
-
Proven track record designing and operating large-scale distributed systems in production
-
Experience setting technical direction and influencing cross-team architecture
-
Advanced proficiency in Go and distributed systems design and cloud-native development
-
Deep expertise in Slurm/Kubernetes internals and cloud-native development
-
Bachelor’s or Master’s degree in CS, EE, or related field
-
Deep expertise in Slurm/Kubernetes internals
-
Familiarity with orchestration and workflow technologies such as Ray, Kubeflow, Kueue, Istio, Knative, or Argo Workflows
-
Experience with distributed workloads, GPU-based applications, or ML pipelines
-
Knowledge of scheduling concepts like quota enforcement, pre-emption, and scaling strategies
-
Exposure to reliability practices including SLOs, alarms, and post-incident reviews
-
Experience with AI infrastructure and workloads (ML training, inference, or HPC)
-
Ability to mentor senior engineers and elevate organizational standards
-
You thrive on solving problems that balance cost, performance, and reliability in high-demand environments
-
You’re curious about orchestration beyond SUNK and how to evolve it for next-generation AI
-
You love defining long-term architecture for systems at global scale
-
You’re an expert at mentorship, architecture, and operational excellence across teams