CoreWeave logo

Staff Software Engineer (Observability)

CoreWeave

On-siteNew York City, NYlead$188k–$300kPosted 7h ago

Job description

  • We are seeking a highly experienced Staff Software Engineer to lead our efforts in building, maintaining, and optimizing highly scalable, reliable, and secure systems

  • The Observability team is responsible for deploying and maintaining critical infrastructure at CoreWeave including our logging, tracing, and metrics platforms as well as the pipelines that feed them

  • You’ll lead and mentor engineers, fostering a culture of collaboration and continuous improvement

  • Scale logging, tracing, and metrics platforms to support a global datacenter footprint

  • Develop and refine monitoring and alerting to enhance system reliability

  • Advise engineers across CoreWeave on optimal usage of Observability systems

  • Automate interactions with CoreWeave’s Compute Infrastructure layer

  • Manage production clusters and ensure development teams follow best practices for deployments- Excellent problem-solving, analytical, and communication skills

  • Expertise in Kubernetes, containerization, and microservices architectures

  • Deep expertise across all observability pillars using tools like ClickHouse, Elastic, Loki, Victoria Metrics, Prometheus, Thanos and/or Grafana

  • 7+ years of experience in Software Engineering, Site Reliability Engineering, DevOps, or a related field

  • Preferred Qualifications:

  • Proven track record of leading incident management and post-mortem analysis

  • Experience running and scaling observability tools as a cloud provider

  • Experience administering large-scale kubernetes clusters

  • Deep understanding of data-streaming systems- Some roles at CoreWeave require a live paired coding exercise in Python or Go