Voxel logo

Senior Software Engineer (ML Infrastructure)

Voxel

On-siteSan Francisco, California, United States {{REMOTE}}senior$200k–$250kPosted 8h ago

Job description

  • We’re hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models

  • You’ll build systems that let our applied ML team train multiple models concurrently, manage experiments and ship optimized models to production

  • You’ll set technical direction, write code, make architecture calls, and partner closely with applied CV, ML Data and Platform engineers

  • Build and maintain training infrastructure that lets the applied ML team train multiple models concurrently, manage experiments, and iterate quickly on new architectures

  • Own the train-to-deploy handoff - export trained models to optimized inference formats (TensorRT, ONNX), quantify accuracy and latency impact, and partner with Platform on production deployment

  • Establish ML experiment tracking and lifecycle management - pick the right tools (Weights & Biases, MLflow, ClearML, or similar) so researchers can run, compare, and reproduce experiments efficiently

  • Establish DevOps-for-ML best practices on AWS (IaC, CI/CD, observability, cost monitoring) so researchers can iterate quickly and safely

  • Understand the infra needs of applied ML/CV engineers and design scalable solutions that support model development

Benefits

  • Extensive / Generous health, dental, and vision insurance

  • Highly competitive paid parental leave and support system

  • Ownership in the business through an Equity Incentive Plan

  • Generous paid time off and / or flexible work arrangements

  • Daily meals in-office, vibrant company events, team-building

  • 401K retirement plan, HSA options, pre-tax Commuter Card- Bias toward shipping. You’d rather ship something good this week than something perfect next quarter

  • 4+ years of experience building and shipping large scale software solutions

  • Track record of owning infrastructure end-to-end: scoping, building, shipping, and improving systems that internal teams depend on

  • Strong communication skills

  • Hands-on experience building ML training pipelines in PyTorch

  • Experience with AWS (S3, EC2, EKS, or similar) for ML workloads

  • Strong Python. Write performant code that scales well in production environments

  • Hands-on experience with ML experiment tracking and lifecycle tools (Weights & Biases, MLflow, ClearML, or similar)

  • Familiarity with GPU performance profiling and optimization (Nsight, PyTorch profiler, or similar)

  • Background in computer vision model training

  • Experience with modern ML orchestration tools (Ray, Sematic, Flyte, Metaflow, Prefect, or similar)