research

MC-PhysBench: A Time-to-Failure Physics Benchmark for World Models

We give video world models one frame of a simple physics scene, ask them to predict what happens next, measure how long each prediction stays physically valid, and document what causes failure.


Loading…

Key Takeaways

6,653
model videos scored
17
models evaluated
9
scenes ranked
50
runs per model, per scene
  • Video world models look physically correct but often aren't. Even on a ball rolling down a ramp, every model breaks physics within a fraction of a second.
  • A newer model generation is not uniformly better. Cosmos 3 improved on Cosmos-Predict2 in some scenes and regressed in others. Wan 2.2 beat Wan 2.1 on one scene and matched it on the rest. Neither came close to ground truth.
  • Simple baselines are difficult to beat. A no-model baseline that just keeps each object moving at its last speed outlasts the models in 91 of 103 results on the rigid scenes.
  • A model can lead on one scene and fail at once on another, which is why every scene is ranked separately and the overall score is only a rough guide.
  • Almost every episode ends in one of two ways: the object moves incorrectly, or the camera moves. Each model tends to fail the same way most of the time.
  • Video input fixes camera control but the physics still breaks. Given the full clip instead of one frame, the Cosmos models keep the camera still, but the ball still disappears or moves the wrong way.
  • Benchmarks based on image similarity miss physics failures. Two models with very different image-similarity scores lasted equally long before breaking physics.

Overview

How the Test Works

  1. We built nine simple physics scenes in a simulator, like a ball rolling down a ramp, two balls colliding, or a ball rolling behind a wall.
  2. We render the first second of each scene, then give the model a single still frame from the end of it, plus a one-line description of the scene.
  3. The model continues the video for two more seconds, predicting what happens next.
  4. We run the same scene in the simulator many times to get the range of things that could physically happen, then check the model's video against it frame by frame, looking for seven kinds of mistake, like an object vanishing or the camera moving.
  5. The clock stops at the first mistake that lasts 0.3 seconds. That time is the model's score, averaged over the same 50 runs of each scene.

Here's an example:

Loading…

Background and Design Decisions

Video world models are judged today by a single score, usually how much their output resembles a reference clip. MC-PhysBench instead asks how long a model's prediction stays physically valid and names what breaks first. Each model continues a rendered physical scene from one frame; its pixels are measured back into object state and checked against a calibrated physics reference. The output is a survival curve per model, its time to failure (how long its prediction stays physically valid), and which of seven checks caught each failure, from objects vanishing to the camera moving. Its measurement error is published with the results.

The scenes are simple: a ball rolling down a ramp, two balls colliding, an object passing behind a wall. Most flagship models lose physical validity within a frame or two of taking over from the clip they are given. If models cannot keep a rolling ball on its path, their output cannot be trusted for the harder worlds people want them for. A resemblance score won't show that.

MC-PhysBench measures time instead of resemblance: when each prediction first breaks physics, over 50 identical episodes per model, compared episode by episode (paired tests). Every failure is traced to a physical cause. The measurement is calibrated and its best possible score, ground truth, is published, so measurement error is kept apart from model error. The ranking rules were fixed before the results existed, including a check that each ranking holds when the tolerance changes.


Analysis

Model by Model

  • Gen-4.5 (Runway): holds the camera and fails on motion; first on the ball through a corridor and tied with H3 Max on the ball into a wall. Fits an image-to-video model built to preserve its input frame.
  • Cosmos 3 Super (NVIDIA, open weights): strongest on contact; beats the no-model baseline on the block stack, but moves the camera on the ball behind a wall unless given video.
  • Veo 3.1 (Google): roughly right for longer than most, never precisely right, so its billiards rank depends on the tolerance.
  • H3 Max (MiniMax): strong physics, unsteady camera; 72 percent of the achievable time on the ball into a wall, but two thirds of episodes end with the camera moving.
  • Seedance 2.5 (ByteDance): no dominant weakness; longest on the ball behind a wall and level with the no-model baseline on both multi-body scenes.
  • Grok Imagine 1.5 (xAI): steady, never a standout; shares the top rank on four scenes but clearly wins none.
  • CogVideoX 1.5 (Zhipu, open weights): the strongest open model outside Cosmos; its best scene is billiards, third longest there, though most models share the top rank on that scene.
  • Cosmos 3 Nano (NVIDIA, open weights): scale alone does not explain the physics; newer and more than twice the size of the 7B it replaces, yet worse on four of six scenes.
  • Wan 2.2 (Alibaba, open weights): an open model keeping pace with hosted flagships on a dynamics scene; longest on the two-ball collision.
  • Gemini Omni (Google): moves things, but not where physics puts them; a frozen frame beats it on every scene.
  • HunyuanVideo 1.5 (Tencent, open weights): fails on motion more than two thirds of the time; longest on no scene and never ahead of the no-model baseline.
  • Kandinsky WM 1.0 (open weights): the longest of any model on the rough ramp (usually the hardest scene), but weak overall.
  • HappyHorse 1.0: re-shoots the scene rather than continuing it; 81 percent of episodes end with the camera moving, 90 on the soft ball.
Every model in one table
Model Strongest scene Weakest scene Most common first failure Below the frozen frame Overall
Gen-4.5 Ball through a corridor, 75% Rough ramp, 2% Wrong motion, 82% 4 of 8 scenes 30%
Cosmos 3 Super Block stack, 53% Rough ramp and ball behind a wall, 3% Wrong motion, 70% 6 of 8 21%
Veo 3.1 Ball through a corridor, 45% Rough ramp, 0% Wrong motion, 66% 4 of 8 20%
H3 Max Ball into a wall, 72% Ball behind a wall, 2% Camera moving, 66% 7 of 8 18%
Seedance 2.5 Ball through a corridor, 34% Ball down a ramp, 5% Wrong motion, 47% 3 of 7 17%
Grok Imagine 1.5 Ball through a corridor, 30% Rough ramp, 2% Wrong motion, 54% 4 of 8 16%
CogVideoX 1.5 Billiards, 32% Ball behind a wall and block stack, 3% Wrong motion, 75% 6 of 8 12%
Cosmos 3 Nano Block stack, 46% Ball through a corridor, 1% Wrong motion, 64% 6 of 8 10%
Gemini Omni Ball into a wall, 30% Ball behind a wall, 2% Wrong motion, 66% 8 of 8 10%
Wan 2.2 Two-ball collision, 29% Ball through a corridor, 1% Wrong motion, 57% 5 of 8 10%
HunyuanVideo 1.5 Two-ball collision, 18% Ball behind a wall and ball through a corridor, 3% Wrong motion, 69% 5 of 8 8%
Kandinsky WM 1.0 Rough ramp, 13% Ball into a wall, 4% Wrong motion, 58% 5 of 8 8%
HappyHorse 1.0 Billiards, 18% Rough ramp, ball down a ramp and ball through a corridor, 2% Camera moving, 81% 7 of 8 6%

Are newer models actually improving at physics?

Within one family, the newer generation is not uniformly better. Two families have an earlier generation on record, run on the same scenes under the same protocol: NVIDIA's Cosmos (Predict2 7B and 14B, then Cosmos 3 Nano and Super) and Alibaba's Wan (2.1, then 2.2). The earlier runs predate the camera-movement check, so the comparison below scores only object motion for both generations, on the generated window.

Cosmos 3 Super more than quadruples the time its predecessor held the ball through the corridor. Cosmos 3 Nano, the smaller tier, moved from instant failure to a few hundredths of a second on the ramps and lost ground on the four other scenes; the older 7B lasted five times longer on billiards. Wan 2.2 improved on the collision and stood still elsewhere. Physics is a stated priority for both families; on these scenes a generation of work has moved the numbers in both directions, and nowhere near ground truth.

Where do models last longest?

No model reaches ground truth on any scene. Every model falls clearly short of it on all nine ranked scenes. Predictions last longest on the two corridor scenes and the block stack, where the best time to failure is 2.0 to 2.3 seconds, and shortest on the rough ramp and the ball behind a wall, where none passes 1.4 seconds.

Each scene has a reference band: how far a prediction can stray from the true path and still count as physically valid. Models are ranked at that band, then re-scored with it halved and doubled. If the order of the models holds at all three widths, the ranking is general: it holds whatever the tolerance. If the order changes, the ranking is band-specific: the rank shown is the one at the scene's own band, and each row shows where the model lands at half and double width. A model that is roughly right for a long time but never precisely right ranks low under a strict tolerance and high under a loose one. That pattern is reported as a result too.

Models that can't be told apart share a rank: two neighbours get different ranks only when a comparison on the same episodes separates them, so several models can share rank 1 (details in Method). Each row shows the model's time to failure with its 95% range on a 0 to 3 second track, with ground truth, the no-model baseline and the frozen frame drawn as lines through every row. Select a row to play that model's prediction beside the ground truth. Survival curves for every model are in Method.

Loading…

What breaks first?

  • Most episodes end with wrong motion or a moving camera. Together these end 88 to 100 percent of episodes on each rigid scene. Objects vanishing ends at most 4 percent, and objects passing through each other almost none. For the median model–scene result no episode is still valid at 3 seconds; the best case is H3 Max on the ball-into-a-wall scene at 44 percent.
  • Models fail in consistent ways. 81 percent of HappyHorse 1.0's episodes and 66 percent of H3 Max's end with the camera moving; 82 percent of Gen-4.5's end with wrong motion.

Each check has its own threshold, set by how much physically valid simulations differ from each other. An episode ends at the first check it fails for 0.3 seconds or more, in the order described under Seven failure checks in Method; slow drift is scored on its own. The bar colours match the check table there.

Loading…

Does video input help?

It stops the camera moving, but not the physics breaking. Only the two Cosmos models were run both ways, on the ball-behind-a-wall scene. Given the whole one-second clip instead of a single frame, they no longer move the camera, which ended 98 percent of their episodes with one frame and none with the clip. Cosmos 3 Super's time to failure rises from 1.05 to 1.44 seconds. But the ball is now lost behind the wall or moves the wrong way instead, and both models still fall significantly short of the no-model baseline (1.76 seconds). The two setups are not ranked against each other.

Loading…

Do image-similarity scores catch these failures?

A video can look right and be physically wrong, and the usual scores cannot tell. Most video benchmarks grade a model by how closely its frames resemble a reference clip, with scores such as SSIM or PSNR. We scored the same episodes both ways. In a separate 30-episode run on the ball-behind-a-wall scene, Cosmos 3 Super scores 0.935 on SSIM against Nano's 0.70, a large gap in appearance, yet their times to failure are indistinguishable: 1.34 against 1.32 seconds. SSIM is rewarding Super for keeping the picture intact while the ball moves wrongly, and penalising Nano for repainting the scene; it cannot tell either apart from correct physics, or say how long a prediction stays valid. That is why this benchmark measures time to failure instead.


Evaluate Your Own Model

Any model that can continue a video can be run on MC-PhysBench. It plugs in through one Python function that takes frames and returns frames.

def generate_fn(prefix_frames: np.ndarray, n_frames: int) -> np.ndarray:
    # prefix_frames: (30, 240, 320, 3) uint8, the rendered 1 s clip at 30 fps
    # returns (n_frames, 240, 320, 3) uint8, your model's continuation

What the function does depends on whether your model takes one image or a clip. The repo has a worked example.

Single-image models. Pass your model the last frame of the clip and the scene's text prompt. Every model on the leaderboard got this input, so your results can be ranked with theirs.

Video-conditioned models. Pass your model the clip, and register it under its own name as video-conditioned. It is reported separately and not ranked against single-image models, which see only one frame.

Running the scenes. A model needs 50 episodes per scene to go on the leaderboard, the same 50 every listed model ran. You can run fewer, but below 50, differences between models are rarely statistically significant. You need Python 3.10+ and a Modal account. The tracker that measures each video runs on Modal GPUs, and your model can too.

For each scene you get time to failure with its 95% range, survival curves, and the check that ended each episode.


Method

Video world models are usually judged by how much their output resembles a reference clip. A clip can score well on resemblance and still get the physics wrong. MC-PhysBench treats the model as a physical predictor and asks how long its prediction stays inside the range of outcomes physics allows, and, when it leaves that range, which law it broke.

1. A scene is simulated and rendered. Each scene is a rigid- or soft-body simulation with a nominal initial state. One episode is one draw from a perturbation model around that state: small changes to positions and velocities, scaled by a per-scene factor. The first second of the episode is rendered to a short video prefix at 30 frames per second. The seed that fixes the perturbation is the same for every model, so every model sees the same 50 episodes.

2. The model continues it. The model receives the last frame of the prefix and the scene's fixed text prompt, and produces a continuation. This is single-image conditioning, the same for every model; two video-conditioned arms that see the whole prefix are shown separately and never ranked against the single-image rows. The continuation is resampled to the scoring resolution and frame rate, then everything after the prefix is scored for two more seconds.

3. Pixels are measured back into state. A fixed measurement operator turns the video into object state: a promptable segmentation tracker follows each object through every frame, including through occlusion, and its masks are reconstructed into positions using the scene's known camera and geometry. No language model is anywhere in this path. The same operator runs on the reference videos, never on simulator state, so the reference and the candidate carry the same instrument error.

4. A reference ensemble sets the tolerance. For every episode, the simulator is run many more times from independently perturbed initial conditions, typically 30 and up to 100, and each run is rendered and measured the same way. How much these physically valid continuations differ from each other defines the tolerance: for each check, the threshold is the 99th percentile of reference-versus-reference deviation, so a correct prediction is flagged about one percent of the time by construction. The band's scale is a declared per-scene parameter, and every ranking is re-checked at half and double that width.

5. The first sustained violation ends the episode. A crossing must persist for 0.3 seconds to count, which keeps single-frame tracker flicker from ending an episode. When several checks cross at once the cause is assigned by a fixed precedence, from camera movement (frame invariance) through objects vanishing (existence), soft-body volume (conservation), objects passing through each other (interpenetration), motion (kinematics), and soft-body shape, so that a moved camera is never misreported as a moved object. The episode's result is a time and a cause. An episode that survives the whole window is censored at 3 seconds and counts as valid.

Access. The closed models were run through Runway's API. Tables name the model; its run id (for example runway_veo3.1) shows on hover.

Seven Failure Checks

Every prediction is checked for seven ways physics can break. An object can vanish, pass through another, move in a way physics doesn't allow, or drift slowly off course; the fixed camera can move; and a soft body can change volume or deform wrongly. Each is a separate check with its own threshold, listed in the table below.

The checks run in a fixed order. Each episode ends at the first check it fails, and when several fail at once, the earliest in the order names the cause (camera movement first, then objects vanishing, soft-body volume, objects passing through each other, motion, and soft-body shape), so a moved camera is never reported as a moved object. Slow drift is scored on its own, outside that order.

Loading…

Survival Analysis

Fifty episodes give a set of failure times with censoring, which is the kind of data survival analysis handles. The Kaplan–Meier estimator turns them into a survival curve: the fraction of episodes still physically valid at each instant. Time to failure, reported as RMVT (the restricted mean validity time), is the area under that curve over the 3-second window: the expected number of seconds an episode stays valid. The first second of that window is the given clip and cannot fail, so every displayed share of ground truth and the overall score use the generated window, RMVT minus 1 s over ground truth's RMVT minus 1 s; subtracting the same second from every row changes no ranking. It is always finite and it is the statistic models are ranked on. VI₅₀ is the time by which half the episodes have failed, where a curve crosses 50 percent; it cannot rank most models, because they fail within a frame or two of the prefix ending and VI₅₀ sits at that floor, about 1.03 seconds.

Loading…

Uncertainty comes from a percentile bootstrap over episodes. Because every model saw the same 50 episodes, model-to-model comparisons use a paired bootstrap on the shared seeds, which cancels the episode-level variance and is more sensitive than comparing two independent intervals. A rank is a dense rank over significant separations: adjacent models are compared with a paired bootstrap on the shared seeds, Holm-corrected within the scene, and rows that cannot be separated share a rank.

Three reference rows sit beside the models on every scene. The no-model baseline (constant velocity) and the frozen frame (copy last state) extrapolate the prefix with no model at all; every model is conditioned on one frame, which carries no speed information, so the frozen frame is the information-matched baseline and the no-model baseline is velocity-advantaged. Ground truth (the instrument ceiling) is the true simulated continuation rendered through the same renderer, codec and tracker, so it shows the longest time to failure the measurement itself can award, between 2.6 and 3.0 seconds here, and every model's time to failure is also reported as a fraction of it.

Validation

A check is only useful if it catches real faults and stays quiet on correct motion. Every check was validated against planted defects before any model was scored, and every validation result is published.

Determinism. The simulator reproduces an episode bit for bit from its seed across processes and machines, so the reference ensemble is the same ensemble wherever it is computed.

Planted defects. Each scene has a library of mutants: a simulated continuation with a known fault inserted at a known time. An object that vanishes, teleports, duplicates, freezes mid-flight, or falls under the wrong gravity; for soft bodies, a wrong stiffness, a volume leak, or a frozen deformation. A scene is admitted only when the clean continuation fires at the designed false-positive rate or below, and each fault is caught by the right channel at the right time.

The cost of measuring through video. The same mutants are passed through the full video path, rendered and tracked, rather than read from simulator state. The difference is the cost of measuring through pixels, reported per scene as the instrument ceiling. Where the tracker cannot see a defect, for instance a shape change smaller than its own noise on a soft body, that limit is recorded.

Resolution. Every population is re-scored at half, one, and twice its reference band from cached trajectories. Where a model's time to failure leaves the floor as the band widens is where it becomes rankable, and every row shows its position at all three widths.


Limitations

Each cell is 50 episodes, so confidence intervals overlap and many models tie. Four of the nine scenes rank only at their own tolerance (band-specific). Gemini, Grok and HappyHorse accept no seed. Sora 2 and Kling are absent because they need their own accounts. The 3-second window is also short.

The checks are necessary but not sufficient. A model that fails one lacks behavioral world understanding, but passing them doesn't prove a model has it.

AuthorShare

Follow what we're building.