
MC-PhysBench
OCT 8, 2026
research
Raising Argus from 29.8% to 50.2% semantic F1 on WGO-Bench with gpt-6.1-sol-high.
Argus is Pantheon's open-source pipeline for annotating robot and human manipulation videos. We took Argus as our foundation to build Midcentury Optimized Argus. Later in this post, we explain the changes to its visual evidence, action boundaries and label generation.
On WGO-Bench's 100 videos, Midcentury Optimized reaches 50.19% semantic F1, up from 29.81% for Argus, a 20.38 percentage-point gain. Both use gpt-6.1-sol-high.1
Three clear Galaxea clips show large gains against the human reference: placing a banana, sorting objects and a drawer sequence. Choose a clip and an action to compare the predictions.
Three Galaxea clips compare Argus and Midcentury Optimized on banana placement, object sorting and a red-drawer sequence.
WGO-Bench tests timestamped, completed manipulation actions. We evaluate all 100 clips and 743 reference actions: 50 DROID, 25 Galaxea and 25 HomER episodes.
Segment F1 measures timing; semantic F1 requires both timing and an accepted action label. Predictions match one-to-one at temporal IoU ≥ 0.75. Micro scores pool events across clips; macro scores average clip-level results. Both systems receive the same videos and task instructions, without reference annotations.
Argus turns a video and its task instruction into timestamped visual evidence, then generates the annotation in one main model pass. For this benchmark, it receives video and camera context, without telemetry or reference labels.
Robot footage is sampled every 1.5 seconds; human egocentric footage, every 0.5 seconds. A text-only router chooses robot grid-cell widths of 224 or 448 pixels from the task instruction. Human grid cells use 256 pixels. Frames retain their source timestamps and are never upscaled. Six-column grids show the episode in sequence, with separate first- and last-frame detail views.
The model predicts action labels and intervals together. Our benchmark adapter asks for WGO-style completed events and permits boundaries between sampled frames; Argus is not restricted to its sampling timestamps. Its native parser extracts the annotations, and validation checks ordered, nonoverlapping intervals. There is no separate boundary-refinement or relabelling pass.
We keep Argus's foundation of timestamped visual evidence and structured annotations, but separate three decisions that the original makes together: what happened, when it happened, and how to describe it.
Read the pipeline from left to right: find the events, settle their boundaries, then label them. All three passes use gpt-6.1-sol-high, without human annotations. The diagrams below show how the evidence changes at each step.
Video and task instruction → propose events and coarse times from 4-FPS grids → refine boundaries using local 8-FPS evidence → pack previous/current/next frames into PNG atlases with role mappings → ground labels at fixed intervals.
First, give the model more chances to see a short action. The two rows below cover the same three seconds of robot footage: Argus samples every 1.5 seconds; we sample every 0.25 seconds, or 4 FPS.
Sampling schematic: over the same three seconds of robot footage, Argus samples every 1.5 seconds and Midcentury Optimized every 0.25 seconds.
These frames form four-column grids, with cells up to 448 pixels for both robot and human footage, without upscaling. The first pass looks for completed actions in that evidence, not just the task's intended outcome. It must find the events now: later passes cannot add a missed pickup or release.
Next, zoom in on where the model thinks an action starts or ends. The full episode remains visible at 1 FPS; each predicted endpoint gets a closer look at 8 FPS within ±1.5 seconds.
Boundary schematic: full-episode context at 1 FPS feeds a local 8-FPS window extending 1.5 seconds on either side of a predicted endpoint.
Only the boundary can move. Events and labels stay fixed, and validation checks the new intervals and limits each endpoint's shift. Argus can already interpolate between frames; the difference here is a fresh model pass with denser local evidence, centered on its own prediction rather than the human reference.
Finally, keep the timing fixed and ask what happened inside the interval. The model sees the previous, current and next actions, with up to five distinct frames per role. That context helps separate the current object's destination from what happens afterward.
Label-grounding schematic: previous, current and next action frames are packed into PNG atlases to produce a label for a fixed interval.
We pack those frames into chronological PNG atlases, with timestamps and role mappings, for up to eight target events per request. Shared captures appear once; oversized sheets split rather than shrink their source pixels. The parser requires one label per target, with no new events and no changes to timing.
Semantic F1 improves across all three source families. Dense first-person HomER footage remains the harder case.
For actions under two seconds, temporal recall rises from 25.57% to 43.75%: 77 of 176 reference actions matched, compared with 45 for Argus. More than half of these shortest actions still remain unmatched.
Across the benchmark, correctly timed and labelled matches rise from 244 to 389, while predicted events fall from 894 to 807. Semantic precision rises from 27.29% to 48.20% and recall from 32.84% to 52.36%: the pipeline recovers more reference actions with fewer incorrect predictions.
These results evaluate the complete pipeline. They do not isolate the contribution of each change.
Published semantic end-to-end F1 results on the original 100-clip WGO-Bench include 16.8% for Macrodata's seeded-relabeling pipeline, 28.0% for Perceptron with task instructions and 38.6% for Vidur. Our 50.19% comes from an internal evaluation. We did not rerun these external systems through our evaluator, so the figures provide context rather than a shared leaderboard.
This work builds on Argus. Separately, we are developing another internal temporal action labelling pipeline at Midcentury. It is a distinct effort from the Argus optimization evaluated here.