
Cloud Inference Engineer
Modular
Job description
-
In the Cloud Inference team, we are focused on building end to end distributed LLM inference deployments that are fully vertically integrated with the MAX stack
-
Our goal is to make inference both the fastest and most scalable while also building an easiest platform for deploying and scaling models for enterprises and developers alike
-
If this sounds exciting, we invite you to join our world-leading AI infrastructure team and help drive our industry forward!
-
Build & ship a LLM focused inference platform using best in class inference techniques (disaggregated inference, multi-node deployment of large models, high performance networking, distributed kv-cache management, high throughput batch processing, etc)
-
Push the envelope for operational excellence with request-to-kernel observability, multi-cloud deployments, clever autoscaling, cold-start optimizations, and more
-
Collaborate with our kernels and genAI teams to achieve SOTA application performance by integrating SOTA kernel & serving optimizations with SOTA cluster optimizations
-
Build helm charts, kubernetes operators, and more to make a create simple, effective, maintainable deployments
Benefits
-
A variety of fantastic health benefits (health, dental, vision insurance; life insurance etc) are available
-
A 401k plan with up to 5% match
-
Free tax advice on Carta
-
Generous work-from-home stipend of $1500 to help you improve your home office
-
Unlimited paid time off and flexible work hours- Weβre seeking engineers who are passionate about pushing the boundaries of distributed inference systems and enjoy working at the intersection of large-scale systems and machine learning
-
We are looking for candidates based on their breadth and depth of experience in backend engineering, AI inference, and distributed systems development
-
Creativity and curiosity for solving complex problems, a team-oriented attitude that enables you to work well with others, and alignment with our culture
-
Experience with kubernetes and operating your own services
-
Strongly identifies with our core company cultural values
-
Ability to create durable, reusable software tools and libraries that are leveraged across teams and functions
-
Experience in machine learning technologies and use cases
-
5+ years of experience working in backend engineering
-
Experience working on high scale ML inference infrastructure (traditional AI or genAI)
-
Familiarity with golang
-
Experience with high performance computing / networking