
Senior Software Engineer (Server Fleet Infrastructure)
CoreWeave
Job description
-
At CoreWeave, infrastructure isn’t just a foundation, it’s a product
-
We build scalable, high-performance computing systems that power the largest AI workloads in the world
-
We’re looking for Engineers that thrive at the intersection of software and systems, deploying and managing large scale bare metal compute
-
Within this domain, you’ll design and build software that manages complex infrastructure across globally distributed datacenters
-
Working in Go, Python/Ansible, deep in Linux environments, observability/monitoring stacks, and leveraging technologies like gRPC and Kubernetes CRs/Controllers/Operators
-
Whether you’re automating bare metal, building fleet lifecycle management services, solving multi-layer integration challenges, or observing our globally distributed fleet, your work will be critical to the company’s delivery of reliable and efficient infrastructure
-
You’ll join a high-performing, communicative, and supportive team, solving complex problems at hyperscale, deploying purpose built AI across the world
-
Design and implement solutions to problems of scale for multi-site deployment and management of CoreWeave’s global server hardware fleet
-
Build and maintain backend services and APIs (gRPC/REST) in Go or Python to interact with Kubernetes and other infrastructure systems
-
Develop provisioning services, automation workflows, and fleet management tools that span from bare metal to container orchestration
-
Write and maintain Kubernetes custom controllers and operators to automate infrastructure behavior
-
Design and implement observability solutions for large-scale server monitoring to improve system stability and insight
-
Adapt and extend open source tooling to enhance visibility into system metrics, performance, and health
-
Create test plans, deployment automation, dashboards, alerts, and insights into our fleet operations
-
Resolve integration challenges across the entire infrastructure stack, from data center hardware to orchestration platforms
-
Participate in an on-call rotation- Familiarity with CI/CD tools like Argo, Flux, and GitHub Actions
-
Strong understanding of Linux internals
-
Proficiency in Go and/or Python software development
-
5+ years of experience in software or infrastructure engineering
-
Experience designing, implementing, and monitoring Kubernetes operators for custom resource definitions
-
Experience with infrastructure automation and configuration management tools like Ansible, Puppet, Chef, Salt
-
Experience with distributed cloud computing principles, including testing strategies, observability, error budgets, and fault-tolerant design
-
Experience implementing metrics pipelines, custom alerts, and monitoring strategies
-
Ability to break down complex problems into achievable tasks and collaborate with teammates to execute them
-
Willingness and ability to thrive in a fast-paced startup environment