CoreWeave logo

Software Engineer (Inference AI/ML)

CoreWeave

On-siteSunnyvale, CAentry$92k–$135kPosted 9h ago

Job description

  • Join the Inference team to ship production features that improve latency, reliability, and cost for model serving on our GPU platform. As an IC1, you’ll implement well-scoped changes, learn our operational practices, and grow quickly with mentorship from experienced engineers

  • Implement well-scoped features and fixes in Python/Go/C++ for model-serving services (e.g., Triton, vLLM, TensorRT-LLM, Ray Serve)

  • Write tests, code comments, and short design docs; participate in code reviews

  • Add basic metrics and dashboards; assist with alarms and runbooks

  • Follow on-call runbooks and learn incident response in a guided rotation

  • Contribute to performance experiments (e.g., request batching, concurrency, caching) with guidance- Foundations in data structures, algorithms, and networked services

  • BS/MS in CS, EE, or related field, or equivalent practical experience

  • Exposure to containers and Kubernetes (coursework or projects welcome)

  • Experience with Python or Go (C++ a plus) and Linux fundamentals; Git/CI basics

  • Curiosity about GPU inference concepts (micro-batching, KV cache, streaming)

  • Coursework/research with PyTorch or TensorFlow; simple CUDA projects a plus

  • Internship or project that deployed a microservice or ML inference demo

  • Familiarity with Grafana/Prometheus/OpenTelemetry or similar tooling