One of the largest egocentric datasets focused on real-world manual labor: roughly 1.9 million hours of first-person video, captured in live work environments and annotated with 3D pose and semantics for Embodied AI and world-model training.
We've built one of the largest egocentric vision datasets focused on real-world manual labor: roughly 1.9M hours of first-person video, collected with custom head-mounted devices in live work environments. Participants wear the device through normal workflows (typically 5–20 hours each), so behavior stays natural across long sequences with minimal inductive bias. These are economically useful physical tasks, captured without scripting or artificial constraints.
It spans light-manufacturing assembly, warehouse logistics, construction and skilled trades, commercial cleaning, food service, and agriculture. The interaction data is dense and high-signal, and hard to replicate synthetically. Footage is 1080p with a wide ~180° field of view, paired with IMU and audio, and hands are visible 80–90% of the time, which suits manipulation-heavy modeling.
On top of the raw data we provide structured annotations and derived signals: 3D pose (with hands), point tracking, motion features, depth where available, and semantic labels covering tasks, environments, actions, objects/tools, and step structure. Everything is filtered with vision-language models for consistency and quality.
Web video is abundant but shallow, with no first-person, contact-rich signal. Real-robot teleoperation captures that signal but can't scale past lab hardware. Egocentric human data keeps the scale of human video and the manipulation fidelity robots need to learn.
Task and difficulty distribution across a representative sample of the collection.
Manufacturing/factory categories are broken out using the same 7 sub-categories and shares as the assembly-manufacturing 50K distribution, so the two line up directly.
Our post-processing turns raw first-person video into structured, learning-ready labels. Three annotation types, at production scale.
Filtered with vision-language models for consistency and quality; frame-accurate segmentation and language descriptions align visual observation, physical motion, and task semantics.
Uniform capture parameters and metadata across the collection, so recordings are directly comparable and trainable out of the box.
Modalities, licensing, task coverage, annotation formats, or a custom collection. Send a question and we'll reply by email.