Cohere logo

Staff Software Engineer (Inference Infrastructure)

Cohere

RemoteleadPosted 9h ago

Job description

  • We are looking for Members of Technical Staff to join the Model Serving team at Cohere

  • The team is responsible for developing, deploying, and operating the AI platform delivering Cohere’s large language models through easy to use API endpoints

  • In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments

  • You will also get the opportunity to interface with customers and create customized deployments to meet their specific needs

Benefits

  • Six weeks’ paid vacation

  • Equity / stock options

  • RRSP, 401(k), and Pension Scheme contributions

  • Coverage for 100% of your insurance premiums across health, dental, vision, and travel

  • Additional coverage for accessing mental health providers/services

  • Six months of fully paid parental leave, including adoption and surrogacy

  • Financial support for egg freezing and IVF in Canada and the UK

  • A monthly fitness and wellness allowance

  • Globally dispersed company that supports a remote work culture

  • A $2,000 annual education benefit for professional development

  • A weekly stipend for meals when working remotely and catered lunch when working from one of our global offices

  • A monthly arts and culture allowance

  • A monthly quality time allowance- Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and build the next generation of AI platforms powering advanced NLP applications?

  • Strong understanding or working experience with distributed systems

  • 5+ years of engineering experience running production infrastructure at a large scale

  • Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving

  • Excellent collaboration and troubleshooting skills to build mission-critical systems, and ensure smooth operations and efficient teamwork

  • The grit and adaptability to solve complex technical challenges that evolve day to day

  • Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters

  • Experience in Golang, C++ or other languages designed for high-performance scalable servers)

  • Familiarity with computational characteristics of accelerators (GPUs, TPUs, and/or custom accelerators), especially how they influence latency and throughput of inference

  • Experience in compute/storage/network resource and cost management

  • Experience with Kubernetes dev and production coding and support

  • If some of the above doesn’t line up perfectly with your experience, we still encourage you to apply!

  • Experience in designing, deploying, supporting, and troubleshooting in complex Linux-based computing environments