
Backend Engineer (Multi-cloud Inference Platform)
Modular
Job description
-
In the Cloud Inference team, we are focused on building end to end distributed LLM inference deployments and a repeatable, observable, productive, low toil platform for managing these deployments
-
Our goal is to make inference both the fastest and most scalable while also building an easiest platform for deploying and scaling models for enterprises and developers alike
-
Build the multi-cloud, multi-tenant platform powering Modular’s inference services
-
Build fault-tolerant, low toil services able to make use of resources in a variety of hardware platforms Clouds (Tier 1 Cloud Providers & neoclouds)
-
Push the envelope for operational excellence with request-to-kernel observability, multi-cloud deployments, cold-start optimizations, and more
-
Build helm charts, kubernetes operators, and more to make a create simple, effective, aintainable deployments
Benefits
-
A variety of fantastic health benefits (health, dental, vision insurance; life insurance etc) are available
-
A 401k plan with up to 5% match
-
Free tax advice on Carta
-
Generous work-from-home stipend of $1500 to help you improve your home office
-
Unlimited paid time off and flexible work hours- We’re seeking engineers who are passionate about pushing the boundaries of distributed inference systems and enjoy working at the intersection of large-scale systems and machine learning
-
We are looking for candidates based on their breadth and depth of experience in backend engineering, AI inference, and distributed systems development. If this sounds exciting, we invite you to join our world-leading AI infrastructure team and help drive our industry forward
-
Experience with kubernetes and operating your own services
-
5+ years of experience working in backend engineering
-
A passion for building and operating high performance, low toil, observable systems
-
Creativity and curiosity for solving complex problems, a team-oriented attitude that enables you to work well with others, and alignment with our culture
-
Strongly identifies with our core company cultural values
-
Experience with Cloud Providers (AWS, GCP, Azure, neoclouds)
-
Experience in machine learning technologies and use cases
-
Practical experience implementing & maintaining security in multi-tenancy environments
-
Experience working on high scale ML inference infrastructure (traditional AI or genAI)
-
Familiarity with golang