
Staff Software Engineer (AI Workload Orchestration)
CoreWeave
Job description
-
As a Staff Software Engineer (IC5) on the AI Workload Orchestration Platform team, you will act as a technical leader for CoreWeave’s Kubernetes-native orchestration strategy for AI workloads
-
You will define and evolve the architecture for how AI workloads are admitted, scheduled, and governed across large GPU clusters using frameworks such as Kueue, Volcano, and Ray
-
This platform serves as a strategic complement to SUNK (Slurm on Kubernetes) and underpins both training and inference workloads across the CoreWeave cloud
-
This role requires strong systems thinking, cross-team influence, and a long-term view of platform scalability, reliability, and developer experience
-
Own the technical vision and architecture for major portions of the AI Workload Orchestration Platform
-
Design scalable, reliable orchestration primitives for AI workloads across multiple schedulers and runtimes
-
Lead cross-team architecture reviews and drive alignment across infrastructure, CKS, and managed inference teams
-
Define platform standards for reliability, observability, capacity management, and operational excellence
-
Identify and resolve systemic performance, scalability, and fairness issues across large GPU clusters
-
Mentor senior engineers and grow technical leadership within the organization
-
Represent the platform in technical reviews and influence broader CoreWeave platform strategy- Strong proficiency in Go and experience designing large-scale, long-lived production systems
-
Deep knowledge of Kubernetes internals, scheduling mechanisms, and controller-based architectures
-
8+ years of professional software engineering experience, with deep expertise in distributed systems or cloud platforms
-
Strong operational mindset with experience owning mission-critical systems at scale
-
Proven ability to lead technical initiatives across teams without direct authority
-
Demonstrated experience designing or evolving orchestration, scheduling, or resource-management platforms
-
Contributions to open-source infrastructure or orchestration projects
-
Experience defining and operating SLOs, capacity models, and large-scale reliability improvements
-
Deep understanding of scheduling concepts including fairness, pre-emption, quota management, and multi-tenant isolation
-
Background in AI infrastructure, ML platforms, HPC, or large-scale batch and streaming systems
-
Hands-on experience with Kueue, Volcano, Ray, or similar Kubernetes-native orchestration frameworks