HiPhy

HiPhy Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation

Tahira Kazimi1, Shubhankar Borse2, Munawar Hayat2, Fatih Porikli2, Pinar Yanardag1

1Virginia Tech   ·   2Qualcomm AI Research

NeurIPS 2026
arXiv Soon Code Soon BibTeX

“A sponge is squeezed over a bowl of water while bubbles rise in a fish tank beside it.”

Elastic deformation · fluid dynamics

Wan

✗ No deformation✗ No bubbles

HiPhy

✓ Elastic deformation✓ Water bubbles

“A hammer strikes a nail into wood while a pot of water boils on the stove beside it.”

Impact mechanics · fluid thermodynamics

Wan

✗ No force transfer✗ No boiling

HiPhy

✓ Nail driven into wood✓ Water boiling

“A balloon floats upward while steam rises from a pot on the stove below it.”

Buoyancy · convective phase transition

Wan

✗ No steam✗ Implausible buoyancy

HiPhy

✓ Steam rising✓ Balloon floating

Abstract

Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video. For example, “a balloon floating upward while steam rises from a pot” requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video.

We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment across various benchmarks, with the largest gains on scenes involving multiple physical principles, where competing methods degrade most sharply.

Method

Models aren’t short on visual capability. Given a detailed enough prompt, they render individual principles with striking fidelity, but the failure shows up specifically when several principles must unfold together, along two distinct axes.

  • 1
    Principle Omission.
    Secondary physics gets dropped entirely: the apple falls realistically, but the water it lands in stays undisturbed.
  • 2
    Temporal Shortcutting.
    Unfolding dynamics collapse into a single visual snapshot: a splash rendered as a static spraying effect instead of progressing through impact → crown formation → submersion.

HiPhy’s dual-level design addresses both directly: per-principle supervision targets Principle Omission, while fine-grained sub-stage reward design addresses Temporal Shortcutting. At the foundational level, HiPhy explicitly optimizes the chronological progression of each physical event by breaking it into smaller stages; at the same time, it aligns these stages with global video coherence and semantic alignment objectives, so multi-principle interactions compose naturally.

HiPhy explicitly decomposes complex physical events, such as an apple falling into water, into temporally ordered sub-stages: Falling, Splash on contact, and Submersion.

HiPhy explicitly decomposes complex physical events (e.g., “an apple falling into water”) into temporally ordered sub-stages. This lets the reward recursively score the chronological progression, completeness, and alignment of the underlying physics.

50K
Training prompts spanning co-occurring physical events
MultiPhyBench
New benchmark for multi-principle physical alignment
Local + Global
Dual-level reward: per-principle stages and whole-scene coherence

Results

HiPhy (right of each pair) versus the Wan base model (left), across rigid-body, fluid, and deformation events.

A colorful rubber ball is dropped from a height, bouncing as it contacts the floor.

Wan

HiPhy

A knife skillfully slices an apple.

Wan

HiPhy

A leaf falls from a tree while rain hits a puddle below.

Wan

HiPhy

A glove catching a fast-moving baseball.

Wan

HiPhy

A gymnast performs an aerial somersault at sunrise.

Wan

HiPhy

A skateboard performs jumps while splashing through a large puddle on a street.

Wan

HiPhy

A stick of incense burns while honey drips slowly from a spoon beside it.

Wan

HiPhy

Camera focuses on a toothpaste tube while a hand squeezes a steady stream of toothpaste.

Wan

HiPhy

A metal can is crushed underfoot while a crumpled piece of paper slowly unfolds and partially recovers its shape.

Wan

HiPhy

A glass bottle rolls off a pier, falls into the water, bobs back up, and floats.

Wan

HiPhy

A woman doing a pirouette in an empty dance studio.

Wan

HiPhy

A slow-motion close-up of a chef slicing a tomato.

Wan

HiPhy

A figure skater executes a powerful spin.

Wan

HiPhy

A ballet dancer twirls on the surface of a still lake at sunset.

Wan

HiPhy

Comparisons

HiPhy against Wan, PnP, PhyT2V, and WISA on prompts that require several concurrent physical processes to resolve correctly.

A book slides over the desk while a small rubber ball falls on desk.

Rigid body dynamics + elastic bouncing

HiPhy

Wan

PnP

PhyT2V

WISA

Coffee accepting a gentle pour of milk.

Fluid mixing

HiPhy

Wan

PnP

PhyT2V

WISA

Sauce simmers in a pan on the stove while a hand is ironing a shirt on a table, and through the kitchen window rain is pouring against the window.

Thermal + mechanical dynamics + fluid impact

HiPhy

Wan

PnP

PhyT2V

WISA

A hand beats eggs in a bowl with a whisk, while a pot of water is boiling next to it, a single candle is also burning on the table.

Mechanical emulsification + convective boiling + combustion

HiPhy

Wan

PnP

PhyT2V

WISA

Ablation

HiPhy’s hierarchical alignment objective is a training signal, not an architecture, so it transfers across different video diffusion backbones.

“A coin dances in a spiral on the table.”

Wan backbone

Base

+ HiPhy

Veo backbone

Base

+ HiPhy

BibTeX

@inproceedings{kazimi2026hiphy, title = {HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation}, author = {Kazimi, Tahira and Borse, Shubhankar and Hayat, Munawar and Porikli, Fatih and Yanardag, Pinar}, booktitle = {Advances in Neural Information Processing Systems (NeurIPS)}, year = {2026} }