
Senior Software Engineer (Cluster Orchestration)
CoreWeave
Job description
-
As part of the Cluster Orchestration team, you will play a key role in advancing CoreWeave’s orchestration platform including SUNK (Slurm on Kubernetes) and beyond our Kubernetes-native foundation that powers AI training and inference at scale
-
This is an opportunity to help shape one of the most critical layers of the AI cloud: ensuring workloads run seamlessly, reliably, and efficiently across massive GPU clusters
-
By building the systems that eliminate infrastructure bottlenecks and create new orchestration capabilities, you will directly empower customers to innovate faster and push the boundaries of what’s possible with AI
-
As a Senior Software Engineer I (IC3), you will own multiple services within the orchestration platform
-
You’ll lead design/code reviews, decompose projects into milestones, and drive measurable improvements in reliability and performance
-
You’ll define SLIs/SLOs for your services, strengthen operational practices, and mentor IC1/IC2 engineers
-
Your work will ensure customers see consistent improvements in throughput, latency, and system resilience- Familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry)
-
Strong coding in Go (Python or C++ a plus) with solid CS fundamentals
-
Hands-on experience running Kubernetes at production scale
-
~3–5 years of professional software engineering experience building distributed systems or cloud services
-
Proven ability to improve service reliability and performance using metrics (P95/P99 latency, throughput, error budgets)
-
Exposure to reliability practices including SLOs, alarms, and post-incident reviews
-
Knowledge of scheduling concepts like quota enforcement, pre-emption, and scaling strategies
-
Familiarity with orchestration and workflow technologies such as Ray, Kubeflow, Kueue, Istio, Knative, or Argo Workflows
-
Experience with distributed workloads, GPU-based applications, or ML pipelines