
Staff Software Engineer (MetalDev)
CoreWeave
Job description
-
As a Staff Software Engineer within our Compute Architecture organization, you will help build the software systems that operate the backbone of our large-scale GPU data centers
-
The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems
-
This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, where your work directly improves the reliability, safety, and scalability of real-world infrastructure.\
-
Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure
-
Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations
-
Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure
-
Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly
-
Translate incidents and hardware failure modes into software improvements that make the platform more resilient
-
Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale
-
Provide technical leadership through design reviews, code reviews, architectural guidance, and mentorship
-
Make pragmatic architecture decisions that balance reliability, simplicity, scalability, and operational burden- Skilled in applying a data-driven approach to reliability, optimization, and continuous improvement
-
B.S., M.S., or PhD in Computer Science or related field, or equivalent experience
-
8+ years of software engineering experience with a strong focus on infrastructure, cloud engineering, and distributed databases—particularly within large-scale datacenter and cloud environments
-
Expertise in Go and proven experience building REST/gRPC APIs for mission-critical platforms
-
Excellent communicator able to work effectively with both technical and non-technical stakeholders
-
Track record of leading incident response, postmortems, and driving robust service reliability
-
Proven success in mentoring engineers, leading technical projects, and influencing engineering strategy across teams
-
Hands-on experience with observability stacks (Prometheus, Grafana, PromQL), CI/CD pipelines, and operating large fleets of GPU servers
-
Strong background in architecting and scaling cloud-native Kubernetes infrastructure and distributed services
-
Experience contributing to and collaborating with open source communities
-
Working knowledge of Kafka, ClickHouse and CRDB
-
DMTF, RedFish APIs, and GPU servers