Physis-LangRead the paper

Physis-LangSelf-Evolving Language as a Physical
Representation for Video World Model

Liming Lu†2, Xianzheng Ma†3, Wenkun He†2, Guanqi Zhan†*1, Yilin Zhao1, Junyu Chen1,
Mengyao Xu1, Jiaojiao Fan1, Wenhang Ge1, Yuchao Gu1, Yunze Liu1, Boyi Li1, Zhen Dong1,
Victor Prisacariu3, Ming-Yu Liu1, Song Han1, Han Cai*1

1 NVIDIA   2 MIT   3 University of Oxford † Equal contribution   * Corresponding authors

Language can do more than describe a scene.
It can represent how the physical world evolves.

Physics-IQ Verified

First and second on the public leaderboard.

#1

Physis-Lang · Cosmos3-Super

48.2 ± 1.4

#2

Physis-Lang · Cosmos3-Nano

43.3 ± 1.5

Super improves on the previous leading score of 42.7 by 5.5 percentage points. Nano also surpasses the original Cosmos3-Super by mean score.

Before and after leaderboard comparison: Physis-Lang Cosmos3-Super scores 48.2 ± 1.4 and Cosmos3-Nano scores 43.3 ± 1.5, taking the top two places ahead of Cosmos3 Super, Seedance 2.5, and MiniMax H3.

The Idea

A shared language for physical intelligence.

Visually convincing videos can still break the laws of physics. Physis-Lang makes causes, interactions, governing principles, and effects explicit in language—then improves that language through feedback. The same representation connects data curation, model training, and inference.

Representation

Beyond what happens. Explain why it happens.

Generated butter softening and spreading as it melts

“Butter melts as the temperature rises.”

Conventional caption
01

Cause

Rising temperature supplies heat to the butter.

02

Physical law

Heat drives melting; gravity makes the butter slump.

03

Effect

The solid shrinks as a shallow liquid pool expands.

Framework

One representation. Three connected stages.

A

Evolve the language

A physics-aware critic diagnoses missing or unsupported claims. An agent refines shared captioning guidelines while the captioner stays fixed.

B

Retrieve missing physics

Model failures become physics-domain queries. Language-guided retrieval finds visually diverse videos that address those deficiencies.

C

Train and generate

The evolved language re-captions training videos and expands inference prompts, with scene-specific negative descriptions to discourage implausible dynamics.

Physis-Lang method: caption evolution loop, physics-tag retrieval, and dataset construction using the best guidelines
Framework from the release paper, Figure 2. Click to view at full resolution.

PhysCapBench

A benchmark for evaluating the richness of physical detail in video captions.

246Reviewed physical videos
3,794Human-verified assertions
15.4Assertions per video
PRECISION

Is each claim supported?

Decompose the caption into atomic claims and verify each against the video as correct, incorrect, or uncertain.

RECALL

Is the physical meaning covered?

Check whether the caption explicitly states or entails each complete human-curated physical assertion.

F1 combines precision and recall. PhysCapBench also provides validation for selecting the best evolved guidelines.

Self-Evolution

Better guidelines. Better physical descriptions. Better physical generation.

PhysCapBench F1

78.64 → 87.82

+9.18 pts

PhyGenBench

64.17 → 67.29

+3.12 pts

Release paper, Figure 4. Downstream comparison keeps pretrained Cosmos3-Nano fixed and changes only the inference captions from guideline iterations 1, 4, 8, and 9. Caption F1 is non-monotonic; iteration 9 is selected.

Language-Guided Data Curation

Find the physics the model is missing.

VideoPhy-2 gains by physical category

Improvement after language-guided retrieval

Cloth deformation (167)
+7.19
Fracture mechanics (94)
+7.45
Elasticity (81)
+1.23
Chemical processes (50)
+8.00
Rigid-body motion (33)
+6.06
Contact / collision (29)
+3.45
Soft-body motion (27)
+7.41
Thermal processes (25)
+8.00

Joint-score gain (percentage points)

Release paper, Figure 5. Numbers in parentheses indicate sample counts. A sample may belong to multiple physical categories.

Results

Stronger physical generation across benchmarks and backbones.

Cosmos3-NanoVeo 3.1Physis-Lang (Cosmos3-Nano)

Release paper, Tables 1–4. Scores use a 100-point scale with the same evaluation protocol per benchmark. VideoPhy-2 above uses the Hard split; on the full set, Physis-Lang scores 68.02 vs. Veo 3.1’s 68.87. PhyGenBench and VideoPhy-2 use GPT-5.5 evaluation.

View benchmark scores and backbone gains
Physical generation benchmark scores · higher is better
ModelPhyGenBenchPhysics-IQ
Verified
VideoPhy-2
All / Hard
PhyGround
Cosmos3-Nano61.6740.2360.41 / 48.3165.18
Veo 3.165.6334.9968.87 / 58.4369.24
Physis-Lang (Cosmos3-Nano)71.0443.4168.02 / 62.3669.90

Gains across model families and scales

Wan

Wan2.114B+7.05

Cosmos

Cosmos3-Edge4B+3.24
Cosmos3-Nano16B+6.22
Cosmos3-Super64B+5.02

Release paper, Table 5. Mean gains are positive for every backbone; individual benchmark gains can vary.

Citation

Build on this work.

@misc{lu2026physislang,
  title={Physis-Lang: Self-Evolving Language as a Physical
         Representation for Video World Model},
  author={Liming Lu and Xianzheng Ma and Wenkun He and Guanqi Zhan
          and Yilin Zhao and Junyu Chen and Mengyao Xu and Jiaojiao Fan
          and Wenhang Ge and Yuchao Gu and Yunze Liu and Boyi Li
          and Zhen Dong and Victor Prisacariu and Ming-Yu Liu
          and Song Han and Han Cai},
  year={2026},
  institution={NVIDIA}
}