
Senior Software Engineer (Inference)
CoreWeave
Job description
-
Senior engineers are area owners who lead designs, raise engineering standards, and deliver measurable improvements to latency, throughput, and reliability across multiple services
-
You’ll partner with product, orchestration, and hardware teams to evolve our Kubernetes-native inference platform and meet strict P99 SLAs at scale
-
Lead design reviews and drive architecture within the team; decompose multi-service work into clear milestones
-
Define and own SLIs/SLOs; ensure post-incident actions land and reliability improves release-over-release
-
Implement advanced optimizations (e.g., micro-batch schedulers, speculative decoding, KV-cache reuse) and quantify impact
-
Strengthen incident posture: capacity planning, autoscaling policy, graceful degradation, rollback/traffic-shift strategies
-
Mentor IC1/IC2 engineers; review cross-team designs and elevate coding/testing standards
-
For IC4: own an area spanning multiple services and teams (e.g., request routing & adaptive scheduling, cost-per-token analytics, GPU resource isolation)- Hands-on experience with Kubernetes at production scale, CI/CD, and observability stacks (Prometheus, Grafana, OpenTelemetry)
-
Practical knowledge of inference internals: batching, caching, mixed precision (BF16/FP8), streaming token delivery
-
Strong coding in Python or Go (C++ a plus) and deep familiarity with networked systems and performance
-
Proven track record improving tail latency (P95/P99) and service reliability through metrics-driven work
-
IC3: ~3–5 years; IC4: ~5–8 years industry experience building distributed systems or cloud services
-
This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency
-
Contributions to inference frameworks (vLLM, Triton, TensorRT-LLM, Ray Serve, TorchServe)
-
Leading multi-team initiatives or partnering with customers on mission-critical launches
-
Experience with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies