research

Optimizing Argus for Action Labelling

Raising Argus from 29.8% to 50.2% semantic F1 on WGO-Bench with gpt-6.1-sol-high.

Overview

Argus is Pantheon's open-source pipeline for annotating robot and human manipulation videos. We took Argus as our foundation to build Midcentury Optimized Argus. Later in this post, we explain the changes to its visual evidence, action boundaries and label generation.

On WGO-Bench's 100 videos, Midcentury Optimized reaches 50.19% semantic F1, up from 29.81% for Argus, a 20.38 percentage-point gain. Both use gpt-6.1-sol-high.1

Grouped bars compare Argus and Midcentury Optimized. Semantic F1 micro: 29.81% and 50.19%; segment F1 micro: 36.04% and 57.94%; semantic F1 macro: 46.32% and 68.68%; segment F1 macro: 52.23% and 77.22%; label accuracy on temporal matches: 82.71% and 86.64%.
Higher is better.

Watch the predictions

Three clear Galaxea clips show large gains against the human reference: placing a banana, sorting objects and a drawer sequence. Choose a clip and an action to compare the predictions.

Three Galaxea clips compare Argus and Midcentury Optimized on banana placement, object sorting and a red-drawer sequence.

What the benchmark measures

WGO-Bench tests timestamped, completed manipulation actions. We evaluate all 100 clips and 743 reference actions: 50 DROID, 25 Galaxea and 25 HomER episodes.

Segment F1 measures timing; semantic F1 requires both timing and an accepted action label. Predictions match one-to-one at temporal IoU ≥ 0.75. Micro scores pool events across clips; macro scores average clip-level results. Both systems receive the same videos and task instructions, without reference annotations.

How Argus works

Argus turns a video and its task instruction into timestamped visual evidence, then generates the annotation in one main model pass. For this benchmark, it receives video and camera context, without telemetry or reference labels.

Robot footage is sampled every 1.5 seconds; human egocentric footage, every 0.5 seconds. A text-only router chooses robot grid-cell widths of 224 or 448 pixels from the task instruction. Human grid cells use 256 pixels. Frames retain their source timestamps and are never upscaled. Six-column grids show the episode in sequence, with separate first- and last-frame detail views.

The model predicts action labels and intervals together. Our benchmark adapter asks for WGO-style completed events and permits boundaries between sampled frames; Argus is not restricted to its sampling timestamps. Its native parser extracts the annotations, and validation checks ordered, nonoverlapping intervals. There is no separate boundary-refinement or relabelling pass.

What we changed

We keep Argus's foundation of timestamped visual evidence and structured annotations, but separate three decisions that the original makes together: what happened, when it happened, and how to describe it.

Read the pipeline from left to right: find the events, settle their boundaries, then label them. All three passes use gpt-6.1-sol-high, without human annotations. The diagrams below show how the evidence changes at each step.

Video and task instruction → propose events and coarse times from 4-FPS grids → refine boundaries using local 8-FPS evidence → pack previous/current/next frames into PNG atlases with role mappings → ground labels at fixed intervals.

Change one: make short events visible

First, give the model more chances to see a short action. The two rows below cover the same three seconds of robot footage: Argus samples every 1.5 seconds; we sample every 0.25 seconds, or 4 FPS.

Sampling schematic: over the same three seconds of robot footage, Argus samples every 1.5 seconds and Midcentury Optimized every 0.25 seconds.

These frames form four-column grids, with cells up to 448 pixels for both robot and human footage, without upscaling. The first pass looks for completed actions in that evidence, not just the task's intended outcome. It must find the events now: later passes cannot add a missed pickup or release.

Change two: revisit the boundary

Next, zoom in on where the model thinks an action starts or ends. The full episode remains visible at 1 FPS; each predicted endpoint gets a closer look at 8 FPS within ±1.5 seconds.

Boundary schematic: full-episode context at 1 FPS feeds a local 8-FPS window extending 1.5 seconds on either side of a predicted endpoint.

Only the boundary can move. Events and labels stay fixed, and validation checks the new intervals and limits each endpoint's shift. Argus can already interpolate between frames; the difference here is a fresh model pass with denser local evidence, centered on its own prediction rather than the human reference.

Change three: ground the label after timing

Finally, keep the timing fixed and ask what happened inside the interval. The model sees the previous, current and next actions, with up to five distinct frames per role. That context helps separate the current object's destination from what happens afterward.

Label-grounding schematic: previous, current and next action frames are packed into PNG atlases to produce a label for a fixed interval.

We pack those frames into chronological PNG atlases, with timestamps and role mappings, for up to eight target events per request. Shared captures appear once; oversized sheets split rather than shrink their source pixels. The parser requires one label per target, with no new events and no changes to timing.

Where the improvement appears

Semantic F1 improves across all three source families. Dense first-person HomER footage remains the harder case.

Semantic and segment F1 by source family. Semantic F1 rises from 48.97 to 74.48 percent on DROID, 42.45 to 74.80 on Galaxea, and 21.96 to 37.28 on HomER.
All 100 clips are included: 50 DROID, 25 Galaxea and 25 HomER.

For actions under two seconds, temporal recall rises from 25.57% to 43.75%: 77 of 176 reference actions matched, compared with 45 for Argus. More than half of these shortest actions still remain unmatched.

Temporal recall for Argus and Midcentury Optimized across five action-duration bins, with gains in every bin.
Short actions remain the most difficult to recover. Recall is measured against reference actions in each duration bin.

Across the benchmark, correctly timed and labelled matches rise from 244 to 389, while predicted events fall from 894 to 807. Semantic precision rises from 27.29% to 48.20% and recall from 32.84% to 52.36%: the pipeline recovers more reference actions with fewer incorrect predictions.

Semantic and segment precision and recall improve for Midcentury Optimized compared with Argus.
Precision measures the share of correct predictions; recall measures the share of reference actions recovered.

These results evaluate the complete pipeline. They do not isolate the contribution of each change.

How this compares with published work

Published semantic end-to-end F1 results on the original 100-clip WGO-Bench include 16.8% for Macrodata's seeded-relabeling pipeline, 28.0% for Perceptron with task instructions and 38.6% for Vidur. Our 50.19% comes from an internal evaluation. We did not rerun these external systems through our evaluator, so the figures provide context rather than a shared leaderboard.

Roadmap

This work builds on Argus. Separately, we are developing another internal temporal action labelling pipeline at Midcentury. It is a distinct effort from the Argus optimization evaluated here.

Evaluation note

  1. Internal evaluation by Midcentury, following the published WGO judging protocol, with Gemini 3.5 Flash at provider-default thinking and temperature. Both systems share our evaluator; exact official-evaluator equivalence is unverified. This is a selected 100-clip run on a development-exposed benchmark, not an untouched holdout.
  2. Dataset: WGO-Bench / Macrodata.

Follow what we're building.