CoreWeave logo

Staff Software Engineer (AI Workload Orchestration)

CoreWeave

On-siteSunnyvale, CAlead$188k–$275kPosted 9h ago

Job description

  • As a Staff Software Engineer (IC5) on the AI Workload Orchestration Platform team, you will act as a technical leader for CoreWeave’s Kubernetes-native orchestration strategy for AI workloads

  • You will define and evolve the architecture for how AI workloads are admitted, scheduled, and governed across large GPU clusters using frameworks such as Kueue, Volcano, and Ray

  • This platform serves as a strategic complement to SUNK (Slurm on Kubernetes) and underpins both training and inference workloads across the CoreWeave cloud

  • This role requires strong systems thinking, cross-team influence, and a long-term view of platform scalability, reliability, and developer experience

  • Own the technical vision and architecture for major portions of the AI Workload Orchestration Platform

  • Design scalable, reliable orchestration primitives for AI workloads across multiple schedulers and runtimes

  • Lead cross-team architecture reviews and drive alignment across infrastructure, CKS, and managed inference teams

  • Define platform standards for reliability, observability, capacity management, and operational excellence

  • Identify and resolve systemic performance, scalability, and fairness issues across large GPU clusters

  • Mentor senior engineers and grow technical leadership within the organization

  • Represent the platform in technical reviews and influence broader CoreWeave platform strategy- Strong proficiency in Go and experience designing large-scale, long-lived production systems

  • Deep knowledge of Kubernetes internals, scheduling mechanisms, and controller-based architectures

  • 8+ years of professional software engineering experience, with deep expertise in distributed systems or cloud platforms

  • Strong operational mindset with experience owning mission-critical systems at scale

  • Proven ability to lead technical initiatives across teams without direct authority

  • Demonstrated experience designing or evolving orchestration, scheduling, or resource-management platforms

  • Contributions to open-source infrastructure or orchestration projects

  • Experience defining and operating SLOs, capacity models, and large-scale reliability improvements

  • Deep understanding of scheduling concepts including fairness, pre-emption, quota management, and multi-tenant isolation

  • Background in AI infrastructure, ML platforms, HPC, or large-scale batch and streaming systems

  • Hands-on experience with Kueue, Volcano, Ray, or similar Kubernetes-native orchestration frameworks