
Staff Software Engineer (Inference Infrastructure)
Cohere
Job description
-
We are looking for Members of Technical Staff to join the Model Serving team at Cohere
-
The team is responsible for developing, deploying, and operating the AI platform delivering Cohere’s large language models through easy to use API endpoints
-
In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments
-
You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs
Benefits
-
Six weeks’ paid vacation
-
Equity / stock options
-
RRSP, 401(k), and Pension Scheme contributions
-
Coverage for 100% of your insurance premiums across health, dental, vision, and travel
-
Additional coverage for accessing mental health providers/services
-
Six months of fully paid parental leave, including adoption and surrogacy
-
Financial support for egg freezing and IVF in Canada and the UK
-
A monthly fitness and wellness allowance
-
Globally dispersed company that supports a remote work culture
-
A $2,000 annual education benefit for professional development
-
A weekly stipend for meals when working remotely and catered lunch when working from one of our global offices
-
A monthly arts and culture allowance
-
A monthly quality time allowance- Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications?
-
Strong understanding or working experience with distributed systems
-
5+ years of engineering experience running production infrastructure at a large scale
-
Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving
-
Excellent collaboration and troubleshooting skills to build mission-critical systems, and ensure smooth operations and efficient teamwork
-
The grit and adaptability to solve complex technical challenges that evolve day to day
-
Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters
-
Experience in Golang, C++ or other languages designed for high-performance scalable servers)
-
Familiarity with computational characteristics of accelerators (GPUs, TPUs, and/or custom accelerators), especially how they influence latency and throughput of inference
-
Experience in compute/storage/network resource and cost management
-
Experience with Kubernetes dev and production coding and support
-
If some of the above doesn’t line up perfectly with your experience, we still encourage you to apply!
-
Experience in designing, deploying, supporting, and troubleshooting in complex Linux-based computing environments