
Staff / Senior Machine Learning Engineer (Reinforcement Learning)
Wayve
Job description
-
As a Senior / Staff Machine Learning Engineer in Wayve’s AV Core organisation, you will advance reinforcement learning methods for end-to-end driving models
-
You will identify where learning from reward or feedback can improve beyond behavior cloning, then take promising ideas from design through large-scale experiments, rigorous evaluation, and integration into our best driving models
-
Driving Core team develops the learning methods that turn diverse driving data into robust closed-loop behavior. You will be a technical owner for reinforcement learning within the group, working closely with researchers and engineers across AV Core, Simulation, Evaluation, and Product Engineering. Success means producing measurable improvements in driving behavior
-
Core Model Safety team develops the core model competencies that enable safe, driverless operation. You will lead the technical direction and delivery of a learned emergency trajectory model for low-frequency, high-consequence maneuvers such as evasive steering and emergency braking
-
You will take the programme from problem definition through modelling, evaluation, integration, and evidence for deployment
-
Shape and execute the reinforcement learning roadmap for Driving Core / Core Model Safety, selecting problems and methods against clear behavioral gaps and measurable success criteria
-
Develop and evaluate post-behavior-cloning optimization methods, including offline and off-policy reinforcement learning as well as other reward-guided approaches; design the regularization, data strategy, and diagnostics needed to make policies reliably better
-
Help improve the reward models and related learning signals used to train and evaluate driving policies, working with partner teams to strengthen their quality, scalability, and downstream usefulness
-
Build robust training and experimentation workflows using large-scale driving data; diagnose distribution shift, objective misspecification, optimization instability, and data or evaluation bias
-
Define evidence across offline metrics, open-loop tests, closed-loop simulation, and on-road evaluation, and distinguish genuine policy improvement from benchmark overfitting
-
Productionize successful methods in the shared ML stack, communicate decisions and results clearly, and raise the technical bar through design reviews, code reviews, and mentoring
Benefits
-
Private healthcare: Choose our optional health insurance for comprehensive coverage for you and your family.
-
Paid time off: Paid vacation plus public holidays and additional leave programs, ensuring you have time to unwind.
-
Mental health resources: Through Spill, you can access therapy and mental health support.
-
Community and socials: Join clubs or attend team socials to connect over hobbies, sports, or just for fun.
-
Competitive compensation: Our compensation package includes cash and equity, making you a true partner in our success.
-
Learning and development: Budgets for books, courses, and company-wide training to support your continuous growth.- Senior-level ownership and collaboration: able to lead a substantial technical area, work across research and engineering boundaries, and bring others along through clear written and verbal communication
-
Excellent experimental judgement: able to turn an ambiguous behavioral problem into falsifiable hypotheses, useful metrics, disciplined ablations, and clear technical decisions
-
Hands-on experience with behaviour cloning, reinforcement learning, or related methods
-
Deep understanding of modern reinforcement learning fundamentals, including policy and value learning, off-policy learning, function approximation, distribution shift, and the failure modes of learned objectives
-
Proficiency in Python and PyTorch, with strong software engineering practices and hands-on experience building reliable machine learning training and evaluation systems
-
A strong track record developing and experimentally validating reinforcement learning or closely related sequential decision-making methods on complex, high-dimensional problems
-
Experience with offline reinforcement learning, imitation learning, reward modeling, preference learning, or post-training of large neural policies
-
Experience in autonomous vehicles, robotics, control, or another domain where policies interact with safety-critical physical systems, including an understanding of motion planning, vehicle dynamics, control, or collision avoidance
-
Experience training multimodal, transformer-based, or generative policy models at scale
-
Experience with closed-loop simulation, off-policy evaluation, uncertainty or calibration, and evaluation under rare or shifted conditions
-
Proficiency in C++, CUDA, distributed training, or performance optimization for production machine learning systems