CoreWeave logo

Senior Software Engineer (Server Fleet Infrastructure)

CoreWeave

On-siteSunnyvale, CAsenior$160k–$225kPosted 9h ago

Job description

  • At CoreWeave, infrastructure isn’t just a foundation, it’s a product

  • We build scalable, high-performance computing systems that power the largest AI workloads in the world

  • We’re looking for Engineers that thrive at the intersection of software and systems, deploying and managing large scale bare metal compute

  • Within this domain, you’ll design and build software that manages complex infrastructure across globally distributed datacenters

  • Working in Go, Python/Ansible, deep in Linux environments, observability/monitoring stacks, and leveraging technologies like gRPC and Kubernetes CRs/Controllers/Operators

  • Whether you’re automating bare metal, building fleet lifecycle management services, solving multi-layer integration challenges, or observing our globally distributed fleet, your work will be critical to the company’s delivery of reliable and efficient infrastructure

  • You’ll join a high-performing, communicative, and supportive team, solving complex problems at hyperscale, deploying purpose built AI across the world

  • Design and implement solutions to problems of scale for multi-site deployment and management of CoreWeave’s global server hardware fleet

  • Build and maintain backend services and APIs (gRPC/REST) in Go or Python to interact with Kubernetes and other infrastructure systems

  • Develop provisioning services, automation workflows, and fleet management tools that span from bare metal to container orchestration

  • Write and maintain Kubernetes custom controllers and operators to automate infrastructure behavior

  • Design and implement observability solutions for large-scale server monitoring to improve system stability and insight

  • Adapt and extend open source tooling to enhance visibility into system metrics, performance, and health

  • Create test plans, deployment automation, dashboards, alerts, and insights into our fleet operations

  • Resolve integration challenges across the entire infrastructure stack, from data center hardware to orchestration platforms

  • Participate in an on-call rotation- Familiarity with CI/CD tools like Argo, Flux, and GitHub Actions

  • Strong understanding of Linux internals

  • Proficiency in Go and/or Python software development

  • 5+ years of experience in software or infrastructure engineering

  • Experience designing, implementing, and monitoring Kubernetes operators for custom resource definitions

  • Experience with infrastructure automation and configuration management tools like Ansible, Puppet, Chef, Salt

  • Experience with distributed cloud computing principles, including testing strategies, observability, error budgets, and fault-tolerant design

  • Experience implementing metrics pipelines, custom alerts, and monitoring strategies

  • Ability to break down complex problems into achievable tasks and collaborate with teammates to execute them

  • Willingness and ability to thrive in a fast-paced startup environment