Midcentury · Egocentric Human Data

First-person data,
at the scale robots need.

One of the largest egocentric datasets focused on real-world manual labor: roughly 1.9 million hours of first-person video, captured in live work environments and annotated with 3D pose and semantics for Embodied AI and world-model training.

1.9M
Hours, first-person
~180°
Field of view
1080p
Capture resolution
5–20 h
Per participant
>85%
Hand visibility
3D pose
+ depth · semantics
The dataset

Real labor, captured first-person

We've built one of the largest egocentric vision datasets focused on real-world manual labor: roughly 1.9M hours of first-person video, collected with custom head-mounted devices in live work environments. Participants wear the device through normal workflows (typically 5–20 hours each), so behavior stays natural across long sequences with minimal inductive bias. These are economically useful physical tasks, captured without scripting or artificial constraints.

It spans light-manufacturing assembly, warehouse logistics, construction and skilled trades, commercial cleaning, food service, and agriculture. The interaction data is dense and high-signal, and hard to replicate synthetically. Footage is 1080p with a wide ~180° field of view, paired with IMU and audio, and hands are visible 80–90% of the time, which suits manipulation-heavy modeling.

On top of the raw data we provide structured annotations and derived signals: 3D pose (with hands), point tracking, motion features, depth where available, and semantic labels covering tasks, environments, actions, objects/tools, and step structure. Everything is filtered with vision-language models for consistency and quality.

The data problem

Where egocentric data sits in the pyramid

Web video is abundant but shallow, with no first-person, contact-rich signal. Real-robot teleoperation captures that signal but can't scale past lab hardware. Egocentric human data keeps the scale of human video and the manipulation fidelity robots need to learn.

SCARCE ·HIGH VALUEABUNDANT ·LOW FIDELITYReal-robot teleoperationhigh-value trajectories · can't scaleSynthetic simulationcontrollable · sim-to-real gapWeb & human videomassive scale · no first-person contact signalEgocentrichuman datascale of video +contact-rich fidelity
Composition

What's in the data

Task and difficulty distribution across a representative sample of the collection.

Task share
Top tasks · share of sample
Sewing garments
13.9% · ~264K hrs
Ironing clothes
10.5% · ~199K hrs
Preparing food items
10.5% · ~199K hrs
Washing dishes
10.5% · ~199K hrs
Metalworking
8.6% · ~164K hrs
Automotive bodywork
7.1% · ~134K hrs
Assembling electronics
7.1% · ~134K hrs
Packaging products
5.8% · ~109K hrs
Preparing & serving food
5.0% · ~95K hrs
Cleaning hotel room
4.7% · ~90K hrs
Cleaning kitchen
4.5% · ~85K hrs
Farm chores
4.2% · ~80K hrs
Repairing motorcycle
3.9% · ~75K hrs
Folding clothes
3.9% · ~75K hrs
Category mix
Share of sample by environment category · modeled hours at 1.9M-hour scale, not an audited per-category count
Food & beverage14.5% · ~275K hrs
Mechanical & small-parts assembly10.0% · ~191K hrs
Construction9.5% · ~180K hrs
Cleaning8.2% · ~156K hrs
Lifestyle / home7.9% · ~150K hrs
Hospitality6.2% · ~118K hrs
Repair services6.0% · ~114K hrs
Fashion5.9% · ~113K hrs
Food processing4.8% · ~92K hrs
Creative workshops3.7% · ~71K hrs
Electronics assembly3.4% · ~65K hrs
Agriculture3.4% · ~64K hrs
Beauty & personal care3.1% · ~58K hrs
Machining & metalwork2.5% · ~47K hrs
Retail2.1% · ~40K hrs
Other non-residential2.0% · ~38K hrs
Healthcare1.9% · ~36K hrs
Automotive & repair0.9% · ~18K hrs
Packing & finishing0.9% · ~18K hrs
Quality inspection0.8% · ~15K hrs
Administrative0.7% · ~14K hrs
Garment & textile manufacturing0.5% · ~9K hrs
Sports & recreation0.3% · ~5K hrs
Entertainment0.3% · ~5K hrs
Printing & design0.2% · ~5K hrs
Laboratory0.1% · ~2K hrs
Difficulty breakdown
Share of sample by task difficulty
Easy
6.7%
Medium
64.4%
Hard
28.9%

Manufacturing/factory categories are broken out using the same 7 sub-categories and shares as the assembly-manufacturing 50K distribution, so the two line up directly.

High-quality annotation

Production-scale 3D pose & semantics

Our post-processing turns raw first-person video into structured, learning-ready labels. Three annotation types, at production scale.

3D hand pose
Tracked in 3D with millimetre-level accuracy; stable under self-occlusion and close-range object interaction.
SLAM-based · mm accuracy
3D full-body pose
Full upper-body and hand articulation reconstructed in 3D, aligned to the head-mounted camera frame.
SLAM-based · mm accuracy
Frame-accurate semantic labels
Action segmentation with explicit language: scene context, action segments, objects/tools, and step structure per demonstration.
Frame-accurate · language
Modalities & derived signals
RGB 1080pIMUAudio3D posePoint trackingMotionDepth
Domains
Food & beverageMechanical & small-parts assemblyConstructionCleaningLifestyle / homeHospitalityRepair servicesFashionFood processingCreative workshopsElectronics assemblyAgriculture+14 more
Quality control

Filtered with vision-language models for consistency and quality; frame-accurate segmentation and language descriptions align visual observation, physical motion, and task semantics.

Capture metrics

Every clip, measured

Uniform capture parameters and metadata across the collection, so recordings are directly comparable and trainable out of the box.

3–30 min
Clip length
5–20 h
Session
MCAP / MP4
Format
30 fps
Frame rate
~180°
Field of view
>85%
Hand visibility
mm-accurate
3D pose
IMU + audio
Sync
Ask the team

Questions about the data?

Modalities, licensing, task coverage, annotation formats, or a custom collection. Send a question and we'll reply by email.