
Senior Software Engineer (Observability Insights)
CoreWeave
Job description
-
We are seeking senior engineers to lead our Observability Insights effort, building the product experiences and agentic interfaces that sit on top of our foundational telemetry layer.
-
You will play a pivotal role in enabling CoreWeave and its customers to understand, troubleshoot, and optimize complex AI systems by delivering core building blocks like multi-tenant APIs, managed Grafana experiences, and MCP-based tool servers.
-
You’ll collaborate closely with PMs and engineering leadership to shape the end-to-end observability experience, providing an outsize opportunity to influence how the world interacts with the forefront of Artificial Intelligence
-
Design and execute the development of highly available, multi-tenant APIs that expose telemetry and derived insights in an developer obsessed way
-
Modernize how users interact with data by building agentic experiences, including MCP servers, agentic tools and API gateways that safely expose foundational telemetry
-
Build agentic observability capabilities that will enable agentic workflows for guided debugging, workload optimization, and incident summarization to empower CoreWeavers and customers alike
-
Develop and enforce best practices regarding the health of telemetry data pipelines, specifically focused on correlation primitives and aggregation services for RCA and performance detection
-
Improve the performance, security, reliability, and scalability of insights services including SLO ownership and latency optimization while participating in the team’s on-call rotation
-
Collaborate closely with internal engineering teams, applying a platform-as-a-product mindset to understand their needs and embed observability best practices and custom tooling into their systems
-
Contribute to the overall observability strategy, influencing the direction of our platform- Experienced in agentic applications or LLM-based features, including grounding, tool calling, and operational safety
-
Familiar with observability systems such as ClickHouse, Loki, VictoriaMetrics, Prometheus, and Grafana
-
Strong focus on developer-facing infrastructure, with a customer-obsessed approach to SDKs, CLIs, and APIs
-
6+ years of experience in software or infrastructure engineering building production-grade backend systems and distributed APIs
-
Comfortable writing production code primarily in Go, with the ability to integrate Python components when needed
-
Collaborative experience in agile teams delivering end-to-end telemetry-to-insights pipelines
-
Proficient in reliability engineering, including fault-tolerant design, SLOs, error budgets, and multi-tenant system resilience
-
Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren’t a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talk
-
You love transforming complex telemetry into actionable insights
-
You’re curious about agentic interfaces and the future of AI observability
-
You’re an expert in building scalable, reliable systems that empower developers and customers alike
-
Hands-on experience with logging, tracing, and metrics platforms in production, with deep knowledge of cardinality, indexing, and query optimization
-
Experience operating Kubernetes clusters at scale, especially for AI workloads
-
Experienced in running distributed systems or API services at cloud scale, including event streaming and data pipeline management
-
Familiarity with LLM frameworks, MCP, and agentic tooling (e.g., Langchain, AgentCore)