Wayve logo

Software Engineer (Data Flywheel Platform)

Wayve

HybridSunnyvale, CAseniorPosted 9h ago

Job description

  • This role sits in the AI Platform organization, on the data flywheel that powers every model we ship

  • Applied Scientists and ML Engineers on the team push the frontier on data curation, enrichment, foundation-model evaluation, and the models themselves

  • This role builds the platform underneath all of it: the pipelines, infrastructure, and systems that turn world-scale fleet data into high-signal training data, evaluate and train foundation models, and enable every team to run these workflows themselves

  • As deployment scales, the leverage is enormous: the better the platform, the faster the whole flywheel turns

  • We are hiring a senior Software Engineer to build the platform that powers Wayve’s data flywheel and foundation-model stack

  • This is the engineering counterpart to our Applied Scientist and ML Engineer roles: you build the systems they, and the wider company, depend on

  • It is high-leverage, high-visibility work with a clear path to deep system ownership

  • Build the systems that allow teams to turn world-scale driving data into high-signal training data, and evaluate and train foundation models on it

  • Replace ad-hoc scripts and manual handoffs with self-serve, observable products used across Science, Autonomy, and Evaluation

  • Every model Wayve ships runs on this platform: your work compounds across the entire fleet and roadmap

  • Work shoulder to shoulder with a world-class science and engineering team, with real deployment at global OEM scale (Nissan, Stellantis, Uber)

  • TC3 / TC4 ownership of platform and infrastructure, with room to set technical direction as the platform matures

  • Build and scale the data curation and enrichment pipelines that turn world-scale fleet data into high-signal training data: mining and active-learning loops, running model-based enrichments over billions of rows, and ensuring data quality at scale

  • Build the evaluation infrastructure behind foundation-model progress: harnesses for offline and closed-loop evaluation, metric and benchmark pipelines, and world-model-based evaluation

  • Build and optimize training and serving infrastructure for large pretrained models: distributed training, batched inference, and large-scale model backfills

  • Build the data-platform backbone: distributed data processing (Ray Data, Daft, Spark / Databricks), embedding and vector search (turbopuffer, Milvus), lakehouse formats (Lance, Iceberg), dataset versioning, and the enrichment and annotation catalog

  • Make it self-serve and reliable: turn one-off processes into products that other teams operate themselves, and own testing, observability, and on-call for what you ship

  • Partner closely with Applied Scientists and ML Engineers to take research from prototype to production at scale

Benefits

  • Private healthcare: Choose our optional health insurance for comprehensive coverage for you and your family.

  • Paid time off: Paid vacation plus public holidays and additional leave programs, ensuring you have time to unwind.

  • Mental health resources: Through Spill, you can access therapy and mental health support.

  • Community and socials: Join clubs or attend team socials to connect over hobbies, sports, or just for fun.

  • Competitive compensation: Our compensation package includes cash and equity, making you a true partner in our success.

  • Learning and development: Budgets for books, courses, and company-wide training to support your continuous growth.- Seniority to match the level: takes ambiguous, cross-team problems and drives them to completion, and at TC4 sets technical direction and multiplies the team

  • A track record of shipping and operating production systems that other teams depend on: testing, code review, observability, and on-call

  • Strong CS fundamentals and several years of production experience (roughly 6 or more for TC3, more for TC4), or equivalent; a degree in CS or comparable practical experience

  • Large-scale data and distributed-systems experience: batch and streaming pipelines, workflow orchestration (Flyte, Airflow, Dagster, or similar), and distributed processing (Spark / PySpark, Ray, Databricks, or equivalent)

  • Systems design for scale: reliable, observable, high-throughput data or ML systems, with strong SQL and query and performance optimization

  • Strong production software engineering, especially production Python (services, APIs, large-scale data processing), and comfort owning and extending large codebases

  • If you’re passionate about self-driving cars and think you have what it takes to make a positive impact on the world, we encourage you to apply

  • ML platform / MLOps: model registration, distributed training, and inference or serving optimization

  • Autonomous driving, robotics, or other large-scale sensor-data workflows

  • Embedding and vector search, annotation tooling, or feature and data catalogs

  • Enough exposure to foundation models, world models, or ML evaluation to partner deeply with scientists

  • Kubernetes and modern data / lakehouse stacks (Databricks, Lance, Iceberg)