
Software Engineer (Inference AI/ML)
CoreWeave
Job description
-
Join the Inference team to ship production features that improve latency, reliability, and cost for model serving on our GPU platform. As an IC1, you’ll implement well-scoped changes, learn our operational practices, and grow quickly with mentorship from experienced engineers
-
Implement well-scoped features and fixes in Python/Go/C++ for model-serving services (e.g., Triton, vLLM, TensorRT-LLM, Ray Serve)
-
Write tests, code comments, and short design docs; participate in code reviews
-
Add basic metrics and dashboards; assist with alarms and runbooks
-
Follow on-call runbooks and learn incident response in a guided rotation
-
Contribute to performance experiments (e.g., request batching, concurrency, caching) with guidance- Foundations in data structures, algorithms, and networked services
-
BS/MS in CS, EE, or related field, or equivalent practical experience
-
Exposure to containers and Kubernetes (coursework or projects welcome)
-
Experience with Python or Go (C++ a plus) and Linux fundamentals; Git/CI basics
-
Curiosity about GPU inference concepts (micro-batching, KV cache, streaming)
-
Coursework/research with PyTorch or TensorFlow; simple CUDA projects a plus
-
Internship or project that deployed a microservice or ML inference demo
-
Familiarity with Grafana/Prometheus/OpenTelemetry or similar tooling