IntroducingMC-EgoHands
A frontier 3D harness for human egocentric video.
01 / MOTION LAB
Motion,
made tangible.
Explore hand motion from recorded RGB sequences, alongside interactive 3D reconstructions.3
Loading…
Loading…
02 / BENCHMARK RESULTS
Four cohorts.
One stereo diagnostic.
PA-MPJPE for the four monocular cohorts, plus three DexYCB stereo metrics. Lower is better.
Loading…
FOUR-COHORT PA OVERVIEW
These overview and distribution figures use PA-MPJPE only.
Loading…
Loading…
Research preview
Current internal results place MC-EgoHands at the state-of-the-art frontier for absolute camera-space MPJPE. A deep, reproducible benchmark is coming soon alongside other releases from Midcentury Labs.
03 / THE 3D HARNESS
From ego video
to robot learning.
MC-EgoHands turns first-person recordings into structured episodes for robot learning workflows.
Loading…
04 / INTEGRATIONS
Designed to meet
your existing stack.
05 / FEATURES
Every stage.
One harness.
Six capability groups connect source recordings to robot-learning experiments. Explore what MC-EgoHands supports today and what comes next.
Loading…
07 / WHAT COMES NEXT
Toward broader
evaluation.
Fuller dataset evaluations are next on the research roadmap.
From video
to measurement.
The system, the experiments and the evidence behind MC-EgoHands.
Loading…
Egocentric video as motion data
Egocentric video records human activity from the wearer’s point of view: the hands, the objects they reach for and the movement between them. MC-EgoHands is a 3D harness for ego video. It connects first-person recordings to structured hand and camera motion, review, robot retargeting and episode export.
The product accepts MP4, IMU JSON and MCAP inputs, with monocular and calibrated stereo processing paths. Source adapters cover MC-EgoHands, EgoDex and HoloAssist; downstream outputs include LeRobot v3 episodes and supported Panda/LIBERO and Yam position, orientation and gripper target layouts. MC-EgoHands annotations are used in VLA and WAM projects with simulation and real-robot experiments. These integrations describe the product workflow; the measurements below evaluate the disclosed hand-reconstruction cohorts.
Midcentury’s egocentric video collection spans 2,000,000+ Hrs. across industrial and everyday environments. This report examines hand reconstruction on four research datasets and a separately disclosed DexYCB stereo diagnostic.1
From pixels to poses
The pipeline starts with RGB images and camera calibration. It detects the hands, reconstructs their joints and surfaces, and estimates their position relative to the camera. The output connects visible activity to a structured representation of hand motion.
Reference hand poses are used to evaluate the predictions. They do not supply pose, depth or scale corrections during inference. Monocular scale depends on the reconstructed hand geometry.2
Evaluation across four datasets
We measured this pipeline on fixed cohorts from ARCTIC, HOT3D, H2O, and FreiHAND. ARCTIC, HOT3D, and H2O test egocentric hand–object interaction; FreiHAND tests independent single-hand images from non-egocentric views. Each cohort has its own camera, annotations, expected-hand definition, and aggregation rule, documented with the result.
| Dataset | Recording context |
|---|---|
| ARCTIC | Bimanual manipulation / moving egocentric camera |
| HOT3D | Hand–object interaction / Project Aria RGB |
| H2O | Bimanual interaction / egocentric cam4 |
| FreiHAND | Single-hand RGB images / third-person views |
ARCTIC, HOT3D and H2O show hands interacting with objects. FreiHAND adds independent single-hand views. The datasets retain their own camera, annotation and scoring conventions.5
Measuring hand pose
PA-MPJPE measures the distance between reconstructed and reference hand joints after alignment. We align each predicted hand independently for translation, rotation and scale, then average the joint distances in millimetres. Lower values indicate a closer match in pose.
Reconstructed joints → align to reference → mean joint distance
ARCTIC, HOT3D and FreiHAND average image-level scores. H2O averages valid hand instances within its cohort. The full alignment and aggregation definitions accompany the data.2
Results across the four datasets
| Dataset | PA-MPJPE (mm) |
|---|---|
| ARCTIC | 3.666 / 5,000 camera images |
| HOT3D | 3.976 / 5,000 camera images |
| H2O | 5.240 / 9,998 hand instances |
| FreiHAND | 4.302 / 5,000 camera images |
- ARCTIC: PA-MPJPE: 21 joints / aligned pose error.
- HOT3D: PA-MPJPE: 21 joints / aligned pose error.
- H2O: PA-MPJPE: 21 joints / aligned pose error.
- FreiHAND: PA-MPJPE: 21 joints / aligned pose error.
The interactive figures show the benchmark means and the shape of their error distributions.
DexYCB stereo diagnostic
This is a post-hoc, ground-truth-informed, unpartitioned diagnostic. It is not a held-out benchmark result.
5,000 dataset frames.
| Metric | Dataset value |
|---|---|
| PA-MPJPE | 3.872 mm |
| Wrist-relative MPJPE | 6.197 mm |
| Absolute camera-space MPJPE | 8.350 mm |
The next evaluation frontier
Our next roadmap expands toward fuller evaluations of the four datasets. We plan to study performance across complete sequences, recording conditions and hand visibility.
The next stages also include matched comparison protocols, uncertainty estimates that account for related video frames, and direct temporal and downstream evaluations. These are planned studies.1
Replaying the measurements
The results were checked by recomputing joint errors from archived predictions and reference poses, then aggregating the benchmark frames again. Each published mean was checked against its metric definition and source records.4
Reading the visual examples
Motion Lab pairs recordings with their 3D hand reconstructions: fifteen bimanual excerpts from ARCTIC, HOT3D and H2O, six curated stereo excerpts from DexYCB, and five independent FreiHAND images. Each example carries its own coverage and error measurements. The DexYCB windows are best-case illustrations from the disclosed post-hoc selected population.3
The motion view uses smoothed playback. The clip measurements reflect those displayed predictions; the four dataset benchmark results remain unchanged. The static artwork at the top is a separate reconstruction study.6
Notes and references
Notes
- Scope and data scale. The 2,000,000+ Hrs. describe Midcentury’s egocentric video collection, separate from the frames evaluated here; they are not a training or inference volume for this report. The four dataset cohorts are custom evaluations, not official leaderboard submissions or matched comparisons with other methods. The DexYCB result is a separate post-hoc diagnostic and is not a held-out benchmark result. These measurements document a processing pipeline rather than a new model architecture. ARCTIC appears in the reconstruction model’s disclosed training mixture; overlap with the other benchmark images has not been independently checked. The measurements do not establish unseen-dataset generalization. Full-dataset evaluations and downstream studies are future work.
- Metric and statistical conventions. PA-MPJPE uses 21 joints and independently fits translation, a positive uniform scale and a proper rotation for each hand in each frame; reflections are excluded. Euclidean distances are reported in millimetres from coordinates archived in metres. ARCTIC, HOT3D and FreiHAND average image-level errors. H2O pools 9,998 valid hand instances within its cohort. FreiHAND uses the proper-rotation companion evaluator; its official evaluator permits an unrestricted orthogonal transform. The two definitions are not interchangeable. Curves connect empirical counts at 0.1 mm intervals. Displayed means use three decimals; the data retain stored precision. No uncertainty interval, repeated-seed estimate or effective independent sample size was computed. Related frames may be correlated. Alignment removes placement, orientation and scale errors, so this score alone does not validate world tracking, physical contact or robot-policy performance. The separate DexYCB diagnostic reports PA-MPJPE, wrist-relative MPJPE and absolute camera-space MPJPE.
- Motion Lab measurements. The Motion Lab contains fifteen bimanual excerpts from ARCTIC, HOT3D and H2O, six curated stereo excerpts from DexYCB, and five independent FreiHAND images. The DexYCB clips are best-case continuous windows from the post-hoc selected diagnostic population. Their 100% expected-hand coverage means the annotated hand has all 21 predicted joints in every displayed frame; it does not establish full visual visibility. Their wrist-relative, absolute camera-space and PA traces come from the frozen SG(9,2) predictions without concatenation or further smoothing. The other fifteen video excerpts report two-hand coverage and stabilized PA-MPJPE. Missing predictions stay missing. FreiHAND images have no temporal filtering or invented motion. The before/after figure compares an acceleration proxy that includes intentional movement as well as jitter; a lower proxy does not establish better detection, world tracking or physical contact. These changes do not alter the benchmark means.
- Verification and evidence access. Archived prediction and reference joints were replayed and reaggregated by cohort. The replay used a 0.0001 mm arithmetic tolerance. This checks numerical agreement, not measurement uncertainty. The audit did not repeat GPU inference or validate reference annotations. Exact checkpoint byte identity is recorded for ARCTIC; the other records identify the frozen pipeline without a run-specific checkpoint hash in this edition. The H2O reconciliation confirms plain PA-MPJPE for its original benchmark cohort; no missing-hand penalty applies to this cohort, and equivalence to another method’s protocol is not claimed. Public downloads contain aggregates, definitions and integrity identifiers. Underlying images, predictions, annotations and execution environments remain access-restricted and are required for independent reproduction. Checksums identify bytes; they do not establish scientific validity or grant access. Research and publication sign-off are pending.
- Dataset composition and valid outputs. ARCTIC (v1_0; egocentric camera 0). HOT3D (hf-30fe9674782f; Project Aria RGB). H2O (v1.1; egocentric cam4). FreiHAND (v2; released per-image camera). Camera calibration is supplied for all four datasets. Reference annotations, including model-fitted labels, have their own errors and modeling assumptions. Per-dataset inputs and validity definitions remain part of the evaluation record.
- Visual studies and attribution. The hero image is a static reconstruction study outside the four benchmarks. Visualization meshes and video encoding are not measurement inputs. Charts use released aggregate observations and zero-based axes. The viewer has no geometry export or download controls. Dataset and component credits identify prior work and do not imply endorsement. Media availability follows the configured website asset release; the numerical report remains readable without video or WebGL.
Dataset references
- Fan et al. ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation. CVPR 2023. Official project page; evaluated release v1_0.
- Banerjee et al. HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos. CVPR 2025. Official project page and citation, CVPR 2025.
- Kwon et al. H2O: Two Hands Manipulating Objects for First Person Interaction Recognition. ICCV 2021. Official dataset repository; evaluated release v1.1.
- Zimmermann et al. FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape from Single RGB Images. ICCV 2019. Official dataset page; evaluated release v2.
- Chao et al. DexYCB: A Benchmark for Capturing Hand Grasping of Objects. CVPR 2021. Official project page; diagnostic source dataset.
Technology credits & licenses
- Potamias, Zhang, Deng and Zafeiriou. WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild. CVPR 2025. Official repository README; component attribution only, no performance comparison.
- Romero, Tzionas and Black. Embodied Hands: Modeling and Capturing Hands and Bodies Together. SIGGRAPH Asia 2017. Official model project page.
The evaluated pipeline uses WiLoR hand reconstruction and the MANO hand model. These components are prior work; their use does not imply endorsement. Interface components: Aceternity UI / Background Beams, shadcn/ui, Recharts, React and Motion. Typography: Satoshi by Fontshare, supplied through the official Fontshare API in upright weights 300, 400, 500, 700 and 900. Interface licenses / Fontshare free font license / 3D renderer license.
MIDCENTURY / EGOCENTRIC DATA
