Software Engineer, Data Infrastructure & Pipelining
Build AI
Visa & sponsorship
- The posting offers relocation assistance.
Job description
About Build AI
Build AI is the data hyperscaler for Physical AI. We co-design hardware, collection, infrastructure, and research to scale the in-the-wild physical labor dataset by orders of magnitude. We learn from humans doing the real job, in real environments. Inflecting revenue, backed by top-tier investors and staffed by leading engineers, Build is becoming the bottleneck to solving physical labor.
Job Summary
Weβre hiring an engineer to own the path from a camera on a worker to training-ready datasets. You will build the tools and infrastructure that offload, store, transform, and serve in-the-wild collection data β on device, in the cloud, and into research. The job is to make that path fast, reliable, and cheap as we add sites, countries, and hours.
Key Responsibilities
-
Design, build, and operate pipelines that move data from collection devices (hats/mounts) through ingest, storage, and into training systems
-
Implement transmission and storage for every stage: on-device capture, upload, object storage, metadata stores, and training-ready shards
-
Optimize end-to-end throughput, cost, and reliability (bandwidth, compression, batching, retries, storage tiers)
-
Architect storage and compute across cloud (and on-prem if we need it); make health, cost, and drop rate obvious
-
Work with research on new data workflows and with Shenzhen firmware so device output is not a snowflake per SKU
-
Build operator tooling: validation, observability, replay, and recovery when the field is messy
You may be a good fit if you have (Must-have qualifications)
-
Strong software engineering fundamentals; Python and at least one other language
-
Experience building reliable backend or data systems (pipelines, distributed storage/ingest)
-
Experience with Linux and with at least one data store (Postgres, MySQL, Elasticsearch, Redis, or equivalent)
-
Comfort working from device to cloud, not only in a warehouse after the fact
-
Bias toward measuring cost and throughput, not just getting a demo working
Strong candidates may also have experience with (Nice-to-have qualifications)
-
Video, pose, or other large media pipelines
-
Cloud infrastructure (AWS, GCP, Azure), orchestration (Kubernetes, Airflow, Temporal), or IaC (Terraform)
-
Dataset management or annotation tooling
-
You have owned cost and throughput of a production data path
Benefits
-
Medical, dental, and vision packages with generous premium coverage
-
$500 per month credit for waiving medical benefits
-
Housing subsidy of $2k per month for those living within walking distance of the office
-
Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)
-
Various wellness benefits covering fitness, mental health, and more
-
Daily lunch and dinner in our office
-
Unlimited compute budget subject to ROI justification
-
Travel
How we're different
Build believes in the Bitter Lesson (http://www.incompleteideas.net/IncIdeas/BitterLesson.html). We are betting early on learning from real human work at massive scale, and that the economies of scale of collection beat extra sensors and extra fidelity. Our addressable market is all physical labor, unlike many of our competitors.
We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.
Build AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply. Questions: research@build.ai