Optimizing Argus for Action Labelling
OCT 5, 2026
research
We give video world models one frame of a simple physics scene, ask them to predict what happens next, measure how long each prediction stays physically valid, and document what causes failure.
Loading…
Here's an example:
Loading…
Video world models are judged today by a single score, usually how much their output resembles a reference clip. MC-PhysBench instead asks how long a model's prediction stays physically valid and names what breaks first. Each model continues a rendered physical scene from one frame; its pixels are measured back into object state and checked against a calibrated physics reference. The output is a survival curve per model, its time to failure (how long its prediction stays physically valid), and which of seven checks caught each failure, from objects vanishing to the camera moving. Its measurement error is published with the results.
The scenes are simple: a ball rolling down a ramp, two balls colliding, an object passing behind a wall. Most flagship models lose physical validity within a frame or two of taking over from the clip they are given. If models cannot keep a rolling ball on its path, their output cannot be trusted for the harder worlds people want them for. A resemblance score won't show that.
MC-PhysBench measures time instead of resemblance: when each prediction first breaks physics, over 50 identical episodes per model, compared episode by episode (paired tests). Every failure is traced to a physical cause. The measurement is calibrated and its best possible score, ground truth, is published, so measurement error is kept apart from model error. The ranking rules were fixed before the results existed, including a check that each ranking holds when the tolerance changes.
| Model | Strongest scene | Weakest scene | Most common first failure | Below the frozen frame | Overall |
|---|---|---|---|---|---|
| Gen-4.5 | Ball through a corridor, 75% | Rough ramp, 2% | Wrong motion, 82% | 4 of 8 scenes | 30% |
| Cosmos 3 Super | Block stack, 53% | Rough ramp and ball behind a wall, 3% | Wrong motion, 70% | 6 of 8 | 21% |
| Veo 3.1 | Ball through a corridor, 45% | Rough ramp, 0% | Wrong motion, 66% | 4 of 8 | 20% |
| H3 Max | Ball into a wall, 72% | Ball behind a wall, 2% | Camera moving, 66% | 7 of 8 | 18% |
| Seedance 2.5 | Ball through a corridor, 34% | Ball down a ramp, 5% | Wrong motion, 47% | 3 of 7 | 17% |
| Grok Imagine 1.5 | Ball through a corridor, 30% | Rough ramp, 2% | Wrong motion, 54% | 4 of 8 | 16% |
| CogVideoX 1.5 | Billiards, 32% | Ball behind a wall and block stack, 3% | Wrong motion, 75% | 6 of 8 | 12% |
| Cosmos 3 Nano | Block stack, 46% | Ball through a corridor, 1% | Wrong motion, 64% | 6 of 8 | 10% |
| Gemini Omni | Ball into a wall, 30% | Ball behind a wall, 2% | Wrong motion, 66% | 8 of 8 | 10% |
| Wan 2.2 | Two-ball collision, 29% | Ball through a corridor, 1% | Wrong motion, 57% | 5 of 8 | 10% |
| HunyuanVideo 1.5 | Two-ball collision, 18% | Ball behind a wall and ball through a corridor, 3% | Wrong motion, 69% | 5 of 8 | 8% |
| Kandinsky WM 1.0 | Rough ramp, 13% | Ball into a wall, 4% | Wrong motion, 58% | 5 of 8 | 8% |
| HappyHorse 1.0 | Billiards, 18% | Rough ramp, ball down a ramp and ball through a corridor, 2% | Camera moving, 81% | 7 of 8 | 6% |
Within one family, the newer generation is not uniformly better. Two families have an earlier generation on record, run on the same scenes under the same protocol: NVIDIA's Cosmos (Predict2 7B and 14B, then Cosmos 3 Nano and Super) and Alibaba's Wan (2.1, then 2.2). The earlier runs predate the camera-movement check, so the comparison below scores only object motion for both generations, on the generated window.
Cosmos 3 Super more than quadruples the time its predecessor held the ball through the corridor. Cosmos 3 Nano, the smaller tier, moved from instant failure to a few hundredths of a second on the ramps and lost ground on the four other scenes; the older 7B lasted five times longer on billiards. Wan 2.2 improved on the collision and stood still elsewhere. Physics is a stated priority for both families; on these scenes a generation of work has moved the numbers in both directions, and nowhere near ground truth.
No model reaches ground truth on any scene. Every model falls clearly short of it on all nine ranked scenes. Predictions last longest on the two corridor scenes and the block stack, where the best time to failure is 2.0 to 2.3 seconds, and shortest on the rough ramp and the ball behind a wall, where none passes 1.4 seconds.
Each scene has a reference band: how far a prediction can stray from the true path and still count as physically valid. Models are ranked at that band, then re-scored with it halved and doubled. If the order of the models holds at all three widths, the ranking is general: it holds whatever the tolerance. If the order changes, the ranking is band-specific: the rank shown is the one at the scene's own band, and each row shows where the model lands at half and double width. A model that is roughly right for a long time but never precisely right ranks low under a strict tolerance and high under a loose one. That pattern is reported as a result too.
Models that can't be told apart share a rank: two neighbours get different ranks only when a comparison on the same episodes separates them, so several models can share rank 1 (details in Method). Each row shows the model's time to failure with its 95% range on a 0 to 3 second track, with ground truth, the no-model baseline and the frozen frame drawn as lines through every row. Select a row to play that model's prediction beside the ground truth. Survival curves for every model are in Method.
Loading…
Each check has its own threshold, set by how much physically valid simulations differ from each other. An episode ends at the first check it fails for 0.3 seconds or more, in the order described under Seven failure checks in Method; slow drift is scored on its own. The bar colours match the check table there.
Loading…
It stops the camera moving, but not the physics breaking. Only the two Cosmos models were run both ways, on the ball-behind-a-wall scene. Given the whole one-second clip instead of a single frame, they no longer move the camera, which ended 98 percent of their episodes with one frame and none with the clip. Cosmos 3 Super's time to failure rises from 1.05 to 1.44 seconds. But the ball is now lost behind the wall or moves the wrong way instead, and both models still fall significantly short of the no-model baseline (1.76 seconds). The two setups are not ranked against each other.
Loading…
A video can look right and be physically wrong, and the usual scores cannot tell. Most video benchmarks grade a model by how closely its frames resemble a reference clip, with scores such as SSIM or PSNR. We scored the same episodes both ways. In a separate 30-episode run on the ball-behind-a-wall scene, Cosmos 3 Super scores 0.935 on SSIM against Nano's 0.70, a large gap in appearance, yet their times to failure are indistinguishable: 1.34 against 1.32 seconds. SSIM is rewarding Super for keeping the picture intact while the ball moves wrongly, and penalising Nano for repainting the scene; it cannot tell either apart from correct physics, or say how long a prediction stays valid. That is why this benchmark measures time to failure instead.
Any model that can continue a video can be run on MC-PhysBench. It plugs in through one Python function that takes frames and returns frames.
def generate_fn(prefix_frames: np.ndarray, n_frames: int) -> np.ndarray:
# prefix_frames: (30, 240, 320, 3) uint8, the rendered 1 s clip at 30 fps
# returns (n_frames, 240, 320, 3) uint8, your model's continuation
What the function does depends on whether your model takes one image or a clip. The repo has a worked example.
Single-image models. Pass your model the last frame of the clip and the scene's text prompt. Every model on the leaderboard got this input, so your results can be ranked with theirs.
Video-conditioned models. Pass your model the clip, and register it under its own name as video-conditioned. It is reported separately and not ranked against single-image models, which see only one frame.
Running the scenes. A model needs 50 episodes per scene to go on the leaderboard, the same 50 every listed model ran. You can run fewer, but below 50, differences between models are rarely statistically significant. You need Python 3.10+ and a Modal account. The tracker that measures each video runs on Modal GPUs, and your model can too.
For each scene you get time to failure with its 95% range, survival curves, and the check that ended each episode.
Video world models are usually judged by how much their output resembles a reference clip. A clip can score well on resemblance and still get the physics wrong. MC-PhysBench treats the model as a physical predictor and asks how long its prediction stays inside the range of outcomes physics allows, and, when it leaves that range, which law it broke.
1. A scene is simulated and rendered. Each scene is a rigid- or soft-body simulation with a nominal initial state. One episode is one draw from a perturbation model around that state: small changes to positions and velocities, scaled by a per-scene factor. The first second of the episode is rendered to a short video prefix at 30 frames per second. The seed that fixes the perturbation is the same for every model, so every model sees the same 50 episodes.
2. The model continues it. The model receives the last frame of the prefix and the scene's fixed text prompt, and produces a continuation. This is single-image conditioning, the same for every model; two video-conditioned arms that see the whole prefix are shown separately and never ranked against the single-image rows. The continuation is resampled to the scoring resolution and frame rate, then everything after the prefix is scored for two more seconds.
3. Pixels are measured back into state. A fixed measurement operator turns the video into object state: a promptable segmentation tracker follows each object through every frame, including through occlusion, and its masks are reconstructed into positions using the scene's known camera and geometry. No language model is anywhere in this path. The same operator runs on the reference videos, never on simulator state, so the reference and the candidate carry the same instrument error.
4. A reference ensemble sets the tolerance. For every episode, the simulator is run many more times from independently perturbed initial conditions, typically 30 and up to 100, and each run is rendered and measured the same way. How much these physically valid continuations differ from each other defines the tolerance: for each check, the threshold is the 99th percentile of reference-versus-reference deviation, so a correct prediction is flagged about one percent of the time by construction. The band's scale is a declared per-scene parameter, and every ranking is re-checked at half and double that width.
5. The first sustained violation ends the episode. A crossing must persist for 0.3 seconds to count, which keeps single-frame tracker flicker from ending an episode. When several checks cross at once the cause is assigned by a fixed precedence, from camera movement (frame invariance) through objects vanishing (existence), soft-body volume (conservation), objects passing through each other (interpenetration), motion (kinematics), and soft-body shape, so that a moved camera is never misreported as a moved object. The episode's result is a time and a cause. An episode that survives the whole window is censored at 3 seconds and counts as valid.
Access. The closed models were run through Runway's API. Tables name the model; its run id (for example runway_veo3.1) shows on hover.
Every prediction is checked for seven ways physics can break. An object can vanish, pass through another, move in a way physics doesn't allow, or drift slowly off course; the fixed camera can move; and a soft body can change volume or deform wrongly. Each is a separate check with its own threshold, listed in the table below.
The checks run in a fixed order. Each episode ends at the first check it fails, and when several fail at once, the earliest in the order names the cause (camera movement first, then objects vanishing, soft-body volume, objects passing through each other, motion, and soft-body shape), so a moved camera is never reported as a moved object. Slow drift is scored on its own, outside that order.
Loading…
Fifty episodes give a set of failure times with censoring, which is the kind of data survival analysis handles. The Kaplan–Meier estimator turns them into a survival curve: the fraction of episodes still physically valid at each instant. Time to failure, reported as RMVT (the restricted mean validity time), is the area under that curve over the 3-second window: the expected number of seconds an episode stays valid. The first second of that window is the given clip and cannot fail, so every displayed share of ground truth and the overall score use the generated window, RMVT minus 1 s over ground truth's RMVT minus 1 s; subtracting the same second from every row changes no ranking. It is always finite and it is the statistic models are ranked on. VI₅₀ is the time by which half the episodes have failed, where a curve crosses 50 percent; it cannot rank most models, because they fail within a frame or two of the prefix ending and VI₅₀ sits at that floor, about 1.03 seconds.
Loading…
Uncertainty comes from a percentile bootstrap over episodes. Because every model saw the same 50 episodes, model-to-model comparisons use a paired bootstrap on the shared seeds, which cancels the episode-level variance and is more sensitive than comparing two independent intervals. A rank is a dense rank over significant separations: adjacent models are compared with a paired bootstrap on the shared seeds, Holm-corrected within the scene, and rows that cannot be separated share a rank.
Three reference rows sit beside the models on every scene. The no-model baseline (constant velocity) and the frozen frame (copy last state) extrapolate the prefix with no model at all; every model is conditioned on one frame, which carries no speed information, so the frozen frame is the information-matched baseline and the no-model baseline is velocity-advantaged. Ground truth (the instrument ceiling) is the true simulated continuation rendered through the same renderer, codec and tracker, so it shows the longest time to failure the measurement itself can award, between 2.6 and 3.0 seconds here, and every model's time to failure is also reported as a fraction of it.
A check is only useful if it catches real faults and stays quiet on correct motion. Every check was validated against planted defects before any model was scored, and every validation result is published.
Determinism. The simulator reproduces an episode bit for bit from its seed across processes and machines, so the reference ensemble is the same ensemble wherever it is computed.
Planted defects. Each scene has a library of mutants: a simulated continuation with a known fault inserted at a known time. An object that vanishes, teleports, duplicates, freezes mid-flight, or falls under the wrong gravity; for soft bodies, a wrong stiffness, a volume leak, or a frozen deformation. A scene is admitted only when the clean continuation fires at the designed false-positive rate or below, and each fault is caught by the right channel at the right time.
The cost of measuring through video. The same mutants are passed through the full video path, rendered and tracked, rather than read from simulator state. The difference is the cost of measuring through pixels, reported per scene as the instrument ceiling. Where the tracker cannot see a defect, for instance a shape change smaller than its own noise on a soft body, that limit is recorded.
Resolution. Every population is re-scored at half, one, and twice its reference band from cached trajectories. Where a model's time to failure leaves the floor as the band widens is where it becomes rankable, and every row shows its position at all three widths.
Each cell is 50 episodes, so confidence intervals overlap and many models tie. Four of the nine scenes rank only at their own tolerance (band-specific). Gemini, Grok and HappyHorse accept no seed. Sora 2 and Kling are absent because they need their own accounts. The 3-second window is also short.
The checks are necessary but not sufficient. A model that fails one lacks behavioral world understanding, but passing them doesn't prove a model has it.