Data Engineer - Video & Multimodal Training Data (World Models)
Our client builds world model systems that learn how the world actually behaves over time and generate it back, interactively, in real time. Seven model releases since December 2024, including the first real-time synchronised audio-and-video world model and a four-player shared world.Data is fast becoming one of the biggest bottlenecks in building world models. Models are only as good as the data behind them, and getting that data right is one of the hardest and most important problems any research lab has.What you'd ownCurating large-scale multimodal datasets across video, robotics and audio to train flagship modelsThe selection logic: filtering, deduplication, quality scoring, captioning and shot detection. The decisions that set the ceiling on model qualityPlatforms processing millions of hours of content, reliably and at a cost that isn't embarrassingDataset versioning, lineage and reproducibility, which show which data was used and whether you can rebuild itWorking directly with researchers on caption quality and data mix, and closing the loop on which data moved which evalSome days you're tuning infrastructure. Some days you're sitting with a researcher, working out why the model learned the wrong thing.What we're looking forYou've built training corpora for visual or multimodal ML at a scale where you had to write code to judge quality because you couldn't review them yourselfHands-on with video and/or audio — codecs, decoding, frame-level work, keeping modalities in sync. Not media as opaque blobs in object storagePython, Kubernetes in production, Ray, Flyte or equivalent, petabyte-scale object storage, sharded dataset formatsYou've deliberately discarded data and can explain exactly how you decidedWe care more about how you think about data than about your years of experience or your publication record. Data engineering or research background — both work.Also worth a conversation if: you've never touched video, but you've run petabyte-to-exabyte data infrastructure where selection, layout, throughput and cost were the whole problem. That skill transfers. They've hired against it before.Why this seat: There is no incumbent data platform and no VP of Engineering in post. You'd define versioning, lineage and the quality bar rather than inherit someone else's from three years ago. The scope doesn't usually fit in one role: video, audio, robotics, synthetic data from our adversarial RL work, and licensed corpora. At a larger lab, that's five teams. You'd own a slice of one. And we name every contributor on our model release pages, including infrastructure and data people.