
Staff Software Engineer (Observability)
CoreWeave
Job description
-
We are seeking a highly experienced Staff Software Engineer to lead our efforts in building, maintaining, and optimizing highly scalable, reliable, and secure systems
-
The Observability team is responsible for deploying and maintaining critical infrastructure at CoreWeave including our logging, tracing, and metrics platforms as well as the pipelines that feed them
-
You’ll lead and mentor engineers, fostering a culture of collaboration and continuous improvement
-
Scale logging, tracing, and metrics platforms to support a global datacenter footprint
-
Develop and refine monitoring and alerting to enhance system reliability
-
Advise engineers across CoreWeave on optimal usage of Observability systems
-
Automate interactions with CoreWeave’s Compute Infrastructure layer
-
Manage production clusters and ensure development teams follow best practices for deployments- Excellent problem-solving, analytical, and communication skills
-
Expertise in Kubernetes, containerization, and microservices architectures
-
Deep expertise across all observability pillars using tools like ClickHouse, Elastic, Loki, Victoria Metrics, Prometheus, Thanos and/or Grafana
-
7+ years of experience in Software Engineering, Site Reliability Engineering, DevOps, or a related field
-
Preferred Qualifications:
-
Proven track record of leading incident management and post-mortem analysis
-
Experience running and scaling observability tools as a cloud provider
-
Experience administering large-scale kubernetes clusters
-
Deep understanding of data-streaming systems- Some roles at CoreWeave require a live paired coding exercise in Python or Go