{"schemaVersion":"jobsearcher.job.v1","id":"16d5effbbd07f858bc95ecaf","url":"https://jobsearcher.com/jobs/16d5effbbd07f858bc95ecaf","canonicalUrl":"https://jobsearcher.com/jobs/16d5effbbd07f858bc95ecaf","title":"Founding Machine Learning Engineer","description":"Own the video data pipeline that frontier robotics labs train on, from raw capture to validated episodes.\r\nHub is one of the fastest-growing data companies, providing real-world training data to the largest AI and robotics companies. Based in San Francisco and backed by Y Combinator and top VCs, our mission is to advance embodied AGI through real-world data and research.\r\nThe RoleAs a founding ML engineer you own Hub's video data pipeline end to end, from raw capture in the field to a dataset a frontier robotics lab trains on. It is the layer that decides what we are allowed to ship.\r\nWhat You'll OwnPetabytes of video from multi-camera rigs we build ourselves: RGB-D, RGB and IMU, metric depth, global and rolling shutter, hardware sync. You turn raw bundles into validated episodes.\r\nEgocentric video with narration, across languages and environments: speech recognition, translation, chapter segmentation, QA against each customer's taxonomy.\r\nQuality control as an ML problem. You benchmark vision-language models, decide which ones we trust to judge our data, fine-tune and distill our own, and own the eval sets and thresholds behind that call.\r\nScalable QC and QA pipelines that trim, cut and quarantine, so no clip ever ships out of spec.\r\nAnnotation at scale against demanding customer taxonomies, plus hand tracking and fine-grained manipulation. Every human verdict becomes a training label.\r\nThe next modalities: tactile and teleoperation sit in the same problem space.\r\nLive production for several of the top 5 AI companies, at high volume, with hard deadlines and specs.\r\nThe ProfileA top engineering school, 3+ years of applied ML in computer vision or robotics. Less experience is fine for outliers: the bar is what you've built.\r\nYou've run VLMs as judges in production: panel agreement, calibration against human labels, fine-tuning and distillation.\r\nYou've trained and deployed CV or robotics models in production: detection and tracking, pose and hand estimation, depth, visual-inertial odometry, action recognition.\r\nYou know the video stack deeply: codecs and frame timing, multi-view geometry, intrinsics and extrinsics, distortion, temporal alignment across sensors.\r\nExceptional individual achievement: elite rankings at competitions or concours, hackathons won, projects at real scale.\r\nAn active GitHub and Hugging Face: recent contributions, open weights and datasets, reproduced results.\r\nYou follow the literature and can tell what's worth implementing from what's noise.\r\nAgentic engineering as a craft: a custom harness, and a loop for your agents to verify their own work through tests, training runs and evals.\r\nNice to HaveYou can come to our Paris office once it opens.\r\nEgocentric vision, IMUs, MCAP, ROS.\r\nWorld models, video generation, VLAs or robot foundation models.\r\nStackPyTorch, CUDA, multi-GPU training and distributed inference. Quantisation, batching and throughput matter as much as accuracy.\r\nVLMs as judges: vLLM serving open-weight models on H100s, hosted models behind one provider interface, and our own fine-tuned and distilled judges.\r\nMulti-stream HEVC video plus high-rate IMU per recording, hardware-synchronised and calibrated, delivered as MCAP.\r\nPostgres for state, S3 for bytes, Kubernetes for compute, GPU inference at petabyte scale. Cost per hour processed is an engineering target.\r\nEvery threshold is a named constant tied to the customer requirement it comes from. Every quarantine carries a code, evidence and an owner.\r\nWhat We Offer$90,000 to $120,000 yearly salary.\r\nStock options between 0.25% and 0.5%, granted at signature.\r\nYour own GPU budget.\r\nBeing able to come to our Paris office once it opens is a big plus.\r\nDirect work with the founders and with the biggest AI labs.#J-18808-Ljbffr","company":"Techtree","rawCompany":"techtree","city":"Millbrae","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-10-05T01:27:18.005Z","occupations":[{"code":"17-2199.08","title":"Robotics Engineers","slug":"robotics-engineers"},{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Founding Machine Learning Engineer","description":"Own the video data pipeline that frontier robotics labs train on, from raw capture to validated episodes.\r\nHub is one of the fastest-growing data companies, providing real-world training data to the largest AI and robotics companies. Based in San Francisco and backed by Y Combinator and top VCs, our mission is to advance embodied AGI through real-world data and research.\r\nThe RoleAs a founding ML engineer you own Hub's video data pipeline end to end, from raw capture in the field to a dataset a frontier robotics lab trains on. It is the layer that decides what we are allowed to ship.\r\nWhat You'll OwnPetabytes of video from multi-camera rigs we build ourselves: RGB-D, RGB and IMU, metric depth, global and rolling shutter, hardware sync. You turn raw bundles into validated episodes.\r\nEgocentric video with narration, across languages and environments: speech recognition, translation, chapter segmentation, QA against each customer's taxonomy.\r\nQuality control as an ML problem. You benchmark vision-language models, decide which ones we trust to judge our data, fine-tune and distill our own, and own the eval sets and thresholds behind that call.\r\nScalable QC and QA pipelines that trim, cut and quarantine, so no clip ever ships out of spec.\r\nAnnotation at scale against demanding customer taxonomies, plus hand tracking and fine-grained manipulation. Every human verdict becomes a training label.\r\nThe next modalities: tactile and teleoperation sit in the same problem space.\r\nLive production for several of the top 5 AI companies, at high volume, with hard deadlines and specs.\r\nThe ProfileA top engineering school, 3+ years of applied ML in computer vision or robotics. Less experience is fine for outliers: the bar is what you've built.\r\nYou've run VLMs as judges in production: panel agreement, calibration against human labels, fine-tuning and distillation.\r\nYou've trained and deployed CV or robotics models in production: detection and tracking, pose and hand estimation, depth, visual-inertial odometry, action recognition.\r\nYou know the video stack deeply: codecs and frame timing, multi-view geometry, intrinsics and extrinsics, distortion, temporal alignment across sensors.\r\nExceptional individual achievement: elite rankings at competitions or concours, hackathons won, projects at real scale.\r\nAn active GitHub and Hugging Face: recent contributions, open weights and datasets, reproduced results.\r\nYou follow the literature and can tell what's worth implementing from what's noise.\r\nAgentic engineering as a craft: a custom harness, and a loop for your agents to verify their own work through tests, training runs and evals.\r\nNice to HaveYou can come to our Paris office once it opens.\r\nEgocentric vision, IMUs, MCAP, ROS.\r\nWorld models, video generation, VLAs or robot foundation models.\r\nStackPyTorch, CUDA, multi-GPU training and distributed inference. Quantisation, batching and throughput matter as much as accuracy.\r\nVLMs as judges: vLLM serving open-weight models on H100s, hosted models behind one provider interface, and our own fine-tuned and distilled judges.\r\nMulti-stream HEVC video plus high-rate IMU per recording, hardware-synchronised and calibrated, delivered as MCAP.\r\nPostgres for state, S3 for bytes, Kubernetes for compute, GPU inference at petabyte scale. Cost per hour processed is an engineering target.\r\nEvery threshold is a named constant tied to the customer requirement it comes from. Every quarantine carries a code, evidence and an owner.\r\nWhat We Offer$90,000 to $120,000 yearly salary.\r\nStock options between 0.25% and 0.5%, granted at signature.\r\nYour own GPU budget.\r\nBeing able to come to our Paris office once it opens is a big plus.\r\nDirect work with the founders and with the biggest AI labs.#J-18808-Ljbffr","datePosted":"2026-10-05T01:27:18.005Z","dateModified":"2026-10-05T01:27:18.005Z","hiringOrganization":{"@type":"Organization","name":"Techtree","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"16d5effbbd07f858bc95ecaf"},"url":"https://jobsearcher.com/jobs/16d5effbbd07f858bc95ecaf"}}