{"schemaVersion":"jobsearcher.job.v1","id":"22aa34a1e906f4860f11d320","url":"https://jobsearcher.com/jobs/22aa34a1e906f4860f11d320","canonicalUrl":"https://jobsearcher.com/jobs/22aa34a1e906f4860f11d320","title":"Software Engineer, Data Infrastructure & Pipelining","description":"About Build AI\n\nBuild AI is the data hyperscaler for Physical AI. We co-design hardware, collection, infrastructure, and research to scale the in-the-wild physical labor dataset by orders of magnitude. We learn from humans doing the real job, in real environments. Inflecting revenue, backed by top-tier investors and staffed by leading engineers, Build is becoming the bottleneck to solving physical labor.\n\nJob Summary\n\nWe’re hiring an engineer to own the path from a camera on a worker to training-ready datasets. You will build the tools and infrastructure that offload, store, transform, and serve in-the-wild collection data — on device, in the cloud, and into research. The job is to make that path fast, reliable, and cheap as we add sites, countries, and hours.\n\nKey Responsibilities\n\nDesign, build, and operate pipelines that move data from collection devices (hats/mounts) through ingest, storage, and into training systems\n\nImplement transmission and storage for every stage: on-device capture, upload, object storage, metadata stores, and training-ready shards\n\nOptimize end-to-end throughput, cost, and reliability (bandwidth, compression, batching, retries, storage tiers)\n\nArchitect storage and compute across cloud (and on-prem if we need it); make health, cost, and drop rate obvious\n\nWork with research on new data workflows and with Shenzhen firmware so device output is not a snowflake per SKU\n\nBuild operator tooling: validation, observability, replay, and recovery when the field is messy\n\nYou may be a good fit if you have (Must-have qualifications)\n\nStrong software engineering fundamentals; Python and at least one other language\n\nExperience building reliable backend or data systems (pipelines, distributed storage/ingest)\n\nExperience with Linux and with at least one data store (Postgres, MySQL, Elasticsearch, Redis, or equivalent)\n\nComfort working from device to cloud, not only in a warehouse after the fact\n\nBias toward measuring cost and throughput, not just getting a demo working\n\nStrong candidates may also have experience with (Nice-to-have qualifications)\n\nVideo, pose, or other large media pipelines\n\nCloud infrastructure (AWS, GCP, Azure), orchestration (Kubernetes, Airflow, Temporal), or IaC (Terraform)\n\nDataset management or annotation tooling\n\nYou have owned cost and throughput of a production data path\n\nBenefits\n\nMedical, dental, and vision packages with generous premium coverage\n\n$500 per month credit for waiving medical benefits\n\nHousing subsidy of $2k per month for those living within walking distance of the office\n\nRelocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)\n\nVarious wellness benefits covering fitness, mental health, and more\n\nDaily lunch and dinner in our office\n\nUnlimited compute budget subject to ROI justification\n\nTravel\n\nHow we're different\n\nBuild believes in the Bitter Lesson. We are betting early on learning from real human work at massive scale, and that the economies of scale of collection beat extra sensors and extra fidelity. Our addressable market is all physical labor, unlike many of our competitors.\n\nWe are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.\n\nBuild AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply. Questions: research@build.ai\n\nCompensation Range: $180K - $300K","company":"Build Ai","rawCompany":"build ai","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-31T11:32:26.819Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Software Engineer, Data Infrastructure & Pipelining","description":"About Build AI\n\nBuild AI is the data hyperscaler for Physical AI. We co-design hardware, collection, infrastructure, and research to scale the in-the-wild physical labor dataset by orders of magnitude. We learn from humans doing the real job, in real environments. Inflecting revenue, backed by top-tier investors and staffed by leading engineers, Build is becoming the bottleneck to solving physical labor.\n\nJob Summary\n\nWe’re hiring an engineer to own the path from a camera on a worker to training-ready datasets. You will build the tools and infrastructure that offload, store, transform, and serve in-the-wild collection data — on device, in the cloud, and into research. The job is to make that path fast, reliable, and cheap as we add sites, countries, and hours.\n\nKey Responsibilities\n\nDesign, build, and operate pipelines that move data from collection devices (hats/mounts) through ingest, storage, and into training systems\n\nImplement transmission and storage for every stage: on-device capture, upload, object storage, metadata stores, and training-ready shards\n\nOptimize end-to-end throughput, cost, and reliability (bandwidth, compression, batching, retries, storage tiers)\n\nArchitect storage and compute across cloud (and on-prem if we need it); make health, cost, and drop rate obvious\n\nWork with research on new data workflows and with Shenzhen firmware so device output is not a snowflake per SKU\n\nBuild operator tooling: validation, observability, replay, and recovery when the field is messy\n\nYou may be a good fit if you have (Must-have qualifications)\n\nStrong software engineering fundamentals; Python and at least one other language\n\nExperience building reliable backend or data systems (pipelines, distributed storage/ingest)\n\nExperience with Linux and with at least one data store (Postgres, MySQL, Elasticsearch, Redis, or equivalent)\n\nComfort working from device to cloud, not only in a warehouse after the fact\n\nBias toward measuring cost and throughput, not just getting a demo working\n\nStrong candidates may also have experience with (Nice-to-have qualifications)\n\nVideo, pose, or other large media pipelines\n\nCloud infrastructure (AWS, GCP, Azure), orchestration (Kubernetes, Airflow, Temporal), or IaC (Terraform)\n\nDataset management or annotation tooling\n\nYou have owned cost and throughput of a production data path\n\nBenefits\n\nMedical, dental, and vision packages with generous premium coverage\n\n$500 per month credit for waiving medical benefits\n\nHousing subsidy of $2k per month for those living within walking distance of the office\n\nRelocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)\n\nVarious wellness benefits covering fitness, mental health, and more\n\nDaily lunch and dinner in our office\n\nUnlimited compute budget subject to ROI justification\n\nTravel\n\nHow we're different\n\nBuild believes in the Bitter Lesson. We are betting early on learning from real human work at massive scale, and that the economies of scale of collection beat extra sensors and extra fidelity. Our addressable market is all physical labor, unlike many of our competitors.\n\nWe are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.\n\nBuild AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply. Questions: research@build.ai\n\nCompensation Range: $180K - $300K","datePosted":"2026-08-31T11:32:26.819Z","dateModified":"2026-08-31T11:32:26.819Z","hiringOrganization":{"@type":"Organization","name":"Build Ai","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"22aa34a1e906f4860f11d320"},"url":"https://jobsearcher.com/jobs/22aa34a1e906f4860f11d320"}}