
Senior Software Engineer (ML Infrastructure)
Voxel
Job description
-
We’re hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models
-
You’ll build systems that let our applied ML team train multiple models concurrently, manage experiments and ship optimized models to production
-
You’ll set technical direction, write code, make architecture calls, and partner closely with applied CV, ML Data and Platform engineers
-
Build and maintain training infrastructure that lets the applied ML team train multiple models concurrently, manage experiments, and iterate quickly on new architectures
-
Own the train-to-deploy handoff - export trained models to optimized inference formats (TensorRT, ONNX), quantify accuracy and latency impact, and partner with Platform on production deployment
-
Establish ML experiment tracking and lifecycle management - pick the right tools (Weights & Biases, MLflow, ClearML, or similar) so researchers can run, compare, and reproduce experiments efficiently
-
Establish DevOps-for-ML best practices on AWS (IaC, CI/CD, observability, cost monitoring) so researchers can iterate quickly and safely
-
Understand the infra needs of applied ML/CV engineers and design scalable solutions that support model development
Benefits
-
Extensive / Generous health, dental, and vision insurance
-
Highly competitive paid parental leave and support system
-
Ownership in the business through an Equity Incentive Plan
-
Generous paid time off and / or flexible work arrangements
-
Daily meals in-office, vibrant company events, team-building
-
401K retirement plan, HSA options, pre-tax Commuter Card- Bias toward shipping. You’d rather ship something good this week than something perfect next quarter
-
4+ years of experience building and shipping large scale software solutions
-
Track record of owning infrastructure end-to-end: scoping, building, shipping, and improving systems that internal teams depend on
-
Strong communication skills
-
Hands-on experience building ML training pipelines in PyTorch
-
Experience with AWS (S3, EC2, EKS, or similar) for ML workloads
-
Strong Python. Write performant code that scales well in production environments
-
Hands-on experience with ML experiment tracking and lifecycle tools (Weights & Biases, MLflow, ClearML, or similar)
-
Familiarity with GPU performance profiling and optimization (Nsight, PyTorch profiler, or similar)
-
Background in computer vision model training
-
Experience with modern ML orchestration tools (Ray, Sematic, Flyte, Metaflow, Prefect, or similar)