CoreWeave logo

Senior Software Engineer (Cluster Orchestration)

CoreWeave

On-siteSunnyvale, CAsenior$139k–$204kPosted 9h ago

Job description

  • As part of the Cluster Orchestration team, you will play a key role in advancing CoreWeave’s orchestration platform including SUNK (Slurm on Kubernetes) and beyond our Kubernetes-native foundation that powers AI training and inference at scale

  • This is an opportunity to help shape one of the most critical layers of the AI cloud: ensuring workloads run seamlessly, reliably, and efficiently across massive GPU clusters

  • By building the systems that eliminate infrastructure bottlenecks and create new orchestration capabilities, you will directly empower customers to innovate faster and push the boundaries of what’s possible with AI

  • As a Senior Software Engineer I (IC3), you will own multiple services within the orchestration platform

  • You’ll lead design/code reviews, decompose projects into milestones, and drive measurable improvements in reliability and performance

  • You’ll define SLIs/SLOs for your services, strengthen operational practices, and mentor IC1/IC2 engineers

  • Your work will ensure customers see consistent improvements in throughput, latency, and system resilience- Familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry)

  • Strong coding in Go (Python or C++ a plus) with solid CS fundamentals

  • Hands-on experience running Kubernetes at production scale

  • ~3–5 years of professional software engineering experience building distributed systems or cloud services

  • Proven ability to improve service reliability and performance using metrics (P95/P99 latency, throughput, error budgets)

  • Exposure to reliability practices including SLOs, alarms, and post-incident reviews

  • Knowledge of scheduling concepts like quota enforcement, pre-emption, and scaling strategies

  • Familiarity with orchestration and workflow technologies such as Ray, Kubeflow, Kueue, Istio, Knative, or Argo Workflows

  • Experience with distributed workloads, GPU-based applications, or ML pipelines