
Lead Machine Learning Inference Engineer (Advertising)
Roku
Job description
-
The Advertising Performance group focuses on performance for all participants in the Advertising ecosystem - Advertisers, Publishers, and Roku. The systems and solutions span multiple disciplines and technologies to perform real-time multi-objective optimization across distributed systems at large scale and with low latency
-
We use Machine Learning, Reinforcement Learning, AI, Control and Optimization Systems, and Auction Dynamics to solve a large set of complex problems. At the core of this is our Machine Learning and Inference Platform that powers the entire landscape
-
In this role, you will architect, design, and lead the development of a SOTA Inference platform that can handle Advertising-level low latencies, scale, throughput, and availability with optimizations that span across hardware, software, and models
-
Weβre looking for a strong technical leader with deep experience in ML serving, high-performance computing, and industry standard frameworks - someone excited to mentor engineers, innovate at scale, and shape the future of machine learning at Roku
-
Lead the design and development of a SOTA Inference platform
-
Oversee the development of monitoring, observability, and other tooling to ensure system and model performance, reliability, and scalability of online inference services
-
Identify and resolve system inefficiencies, performance bottlenecks, and reliability issues, ensuring optimized end-to-end performance
-
Stay at the forefront of advancements in inference frameworks, ML hardware acceleration, and distributed systems, and incorporate innovations where and when they are impactful
Benefits
-
Medical, wellness and financial benefits
-
Free snacks and access to the company fitness center
-
Unlimited paid time off policy
-
Work from home opportunities- 10+ years of experience in developing and deploying large-scale, distributed systems, with at least 5 years in a leadership or technical lead role
-
Proven experience optimizing performance for large-scale machine learning systems, including a deep knowledge of SOTA model optimizations, hardware-software co-design, GPU acceleration, and HPC techniques
-
Experience leading teams working on high-throughput, low-latency ML serving systems
-
Strong programming skills in high-performance languages
-
Deep understanding of inference frameworks and ML system deployment
-
Experience collaborating with and leading global, cross-functional teams
-
Contributions to open-source ML or systems projects
-
Excellent communication and collaboration skills
-
M.S. or above in CS, ECE, or a related field