
Software Engineer (C++ GPU Performance)
Zoox
Job description
-
As a GPU performance software engineer within the Software Performance team, you will instrument, monitor, analyze and optimize GPU-based algorithms that are performance-critical for our solution
-
The scope for GPU usage ranges from traditional computer vision and deep learning architectures to complex geometric reasoning and multi-agent decision making
-
Your work will strongly influence design decisions of future compute platforms & resource allocation
-
Build real-time instrumentation for performance monitoring (CPU, GPU, latency, memory) and develop offline benchmarking frameworks, tools, and scripts to evaluate & analyze performance at scale in CI/vehicle, and establish budgets for next-gen architectures
-
Analyze performance metrics to identify GPU hotspots and root causes, and propose and co-implement actionable solutions with component teams
-
Support teams on bringing serial algorithms to the GPU to maximize compute utilization and improve overall latency
-
Work as part of the Core team to design a middleware framework that promotes by default efficient and performant code development by maximizing CPU and GPU
Benefits
-
Paid parental leave
-
Affinity groups and sports clubs
-
Work from home opportunities
-
Health insurance
-
Our crew’s health and happiness is our first priority. We offer comprehensive health and mental health support, a wellbeing program, and unlimited and flexible paid time away
-
We invest in our crew—and their families—for the long term. That includes generous family planning support, caregiver support, and strong cash compensation with great equity upside
-
We look after our crew when they’re in the office too. Our famous food program is a great example, featuring a daily changing menu of local and sustainable dishes
-
There’s a busy calendar of social events at Zoox, with more sports teams than you can count. And, of course, playing with robots is an important part of the job description- Strong knowledge of C++ and experience in large code bases, comfortable in Linux development environments
-
Experience in development, debugging, and profiling of complex multiprocess systems (e.g., robotic systems, game engines)
-
BS in computer science or related field and 3+ years of experience
-
Strong knowledge of CUDA as applied to recent GPU microarchitectures (e.g., Ampere, Blackwell) and experience debugging/optimizing GPU kernels using tools like Nsight
-
You do not need to match every listed expectation to apply for this position
-
Proficiency with SQL, DataBricks, Looker, or other business intelligence tools
-
Hands-on work with ML model optimization (post-training quantization, layer pruning, etc) or hand-tuning GPU kernels (in OpenGL, CUDA, RocM or similar)
-
Experience with GPU kernel development in a real-time environment, including PTX-level programming, CPU SIMD instructions (e.g., AVX intrinsics), and custom CUDA layers with frameworks like TensorRT & XLA