Wayve logo

Machine Learning Engineer (Performance Tooling)

Wayve

On-site🇬🇧London, United KingdommidPosted 1d ago

Visa & sponsorship

  • UK Licensed Sponsor (group): matched through a parent or group company on the official sponsor register.

Job description

  • You’ll join the AI Performance Tooling team within Wayve’s AI Performance organization, which makes model training and inference faster, more efficient, and more predictable across cloud and embedded hardware.

  • Our mission is to enable data-driven AI performance decisions across priority workloads and hardware targets: Measure performance teams can trust; Monitor trends and catch regressions; Predict the cost of changes before we run them; Advise on bottlenecks and prioritized opportunities.

  • You’ll build tools that reason across the AI stack — models and operators, compilers, runtimes, accelerators, and distributed training infrastructure — turning profiling data into a clear picture of where time, memory, power, and compute go, and what proposed changes will do to latency, throughput, compute spend, and capacity.

  • You’ll work closely with model, compiler, runtime, platform, and hardware teams, bringing a cross-stack view that turns measurement into clear recommendations

  • Design and build reliable, self-service performance tools that scale across models, hardware targets, and development workflows

  • Shape how Wayve measures and predicts AI performance, and set the standards other teams build on

  • Model theoretical peak for a platform, compare with achieved performance, and pinpoint where efficiency is lost at layer and op level

  • Predict latency, memory, utilization, and compute cost of a model or recipe change before spending compute

  • Own monitoring and regression alerting across model builds and training runs

  • Work with training and runtime engineers to set performance targets and make the case with data

Benefits

  • Private healthcare: Choose our optional health insurance for comprehensive coverage for you and your family.

  • Paid time off: Paid vacation plus public holidays and additional leave programs, ensuring you have time to unwind.

  • Mental health resources: Through Spill, you can access therapy and mental health support.

  • Community and socials: Join clubs or attend team socials to connect over hobbies, sports, or just for fun.

  • Competitive compensation: Our compensation package includes cash and equity, making you a true partner in our success.

  • Learning and development: Budgets for books, courses, and company-wide training to support your continuous growth.- Hands-on experience developing deep learning models with PyTorch

  • Judgment to turn an ambiguous performance question into a measurable one, and to prioritize what matters

  • Deep, hands-on performance engineering in complex systems: profiling, roofline analysis, latency and throughput optimization, and root-causing what limits a workload

  • Quantitative communication clear enough to influence another team’s priorities

  • A track record of owning a tool or service end to end — design, delivery, and adoption by other teams

  • Strong Python skills, and comfort profiling and instrumenting large production codebases

  • Data analysis skills to turn noisy measurements into conclusions you can defend