CoreWeave logo

Staff Software Engineer (MetalDev)

CoreWeave

On-siteSunnyvale, CAlead$207k–$275kPosted 11h ago

Job description

  • As a Staff Software Engineer within our Compute Architecture organization, you will help build the software systems that operate the backbone of our large-scale GPU data centers

  • The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems

  • This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, where your work directly improves the reliability, safety, and scalability of real-world infrastructure.\

  • Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure

  • Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations

  • Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure

  • Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly

  • Translate incidents and hardware failure modes into software improvements that make the platform more resilient

  • Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale

  • Provide technical leadership through design reviews, code reviews, architectural guidance, and mentorship

  • Make pragmatic architecture decisions that balance reliability, simplicity, scalability, and operational burden- Skilled in applying a data-driven approach to reliability, optimization, and continuous improvement

  • B.S., M.S., or PhD in Computer Science or related field, or equivalent experience

  • 8+ years of software engineering experience with a strong focus on infrastructure, cloud engineering, and distributed databases—particularly within large-scale datacenter and cloud environments

  • Expertise in Go and proven experience building REST/gRPC APIs for mission-critical platforms

  • Excellent communicator able to work effectively with both technical and non-technical stakeholders

  • Track record of leading incident response, postmortems, and driving robust service reliability

  • Proven success in mentoring engineers, leading technical projects, and influencing engineering strategy across teams

  • Hands-on experience with observability stacks (Prometheus, Grafana, PromQL), CI/CD pipelines, and operating large fleets of GPU servers

  • Strong background in architecting and scaling cloud-native Kubernetes infrastructure and distributed services

  • Experience contributing to and collaborating with open source communities

  • Working knowledge of Kafka, ClickHouse and CRDB

  • DMTF, RedFish APIs, and GPU servers