MiniMax-H3 Level-of-Token research release

Experimental LoT adapter and rank-16 DiT LoRA, with benchmarks and visual examples.

Research status: the latest campaign found no branch that met the predeclared promotion threshold. Metrics show tradeoffs, not a confirmed quality win. See benchmark results and campaign summary.

What was tried

We measured token-layout speed, screened sparse layouts, trained on still images and OpenVidHD clips, compared a control against dense-teacher preservation, and tested inference-time detail/entropy selectors. The best measured DiT-forward speedup was 2.64× on a 37-frame 768×1344 case; VAE decode was unchanged. The r4 500-update branches were effectively tied on the 24-source confirmation set and neither passed promotion criteria. Selector previews were small validation screens, not confirmation benchmarks. Read the experiment history and values.

Weights

Four-column video comparisons

Both grids show matched 30-step outputs from the untouched confirmation prompts. The fourth column shows the matching token layout for each row (latent t=0); source-reference rows are marked as having no model tile mask. Full seven-frame layouts appear in the prompt sheets below. The downloadable MP4 links are generated outputs.

Control branch, step 500

Control comparison with tile-mask column

Preserve branch, step 500

Preserve comparison with tile-mask column

Prompts, videos, and tile-layout masks

The mask sheets show all seven latent-time layouts for each prompt. Each sheet compares dense 1×1, uniform 1×2, and the adaptive detail50 pattern recorded at the warm-up selection step. These are token-layout maps, not semantic segmentation masks. Each prompt has separate masks for the control and preserve checkpoints.

Prompt p0

In the video, a man wearing glasses and a blue shirt is seen interacting with a small dog. The dog, with its white and brown fur, is sitting on a pink surface. The man appears to be speaking to the dog, possibly giving it a command or a treat. The scene is set against a white wall, which provides a neutral backdrop for the interaction between the man and the dog. The overall style of the video is casual and intimate, capturing a moment of connection between the man and his pet.

Control branch — step 500 Preserve branch — step 500
Control p0 tile masks Preserve p0 tile masks
Branch Dense LoRA video Adaptive detail50 video Uniform 1×2 video
Control MP4 MP4 MP4
Preserve MP4 MP4 MP4

Prompt p1

In the video, a man is seen interacting with a yellow and black device attached to a green door. The device appears to be a camera or a sensor, and the man is holding it up to the door, possibly taking a picture or recording a video. The man is dressed in a black jacket and seems to be focused on his task. The door is green and has a window, through which the man's reflection can be seen. The scene suggests that the man might be a security officer or a technician, inspecting or setting up the device for security purposes. The overall style of the video is realistic and straightforward, with no additional elements or distractions.

Control branch — step 500 Preserve branch — step 500
Control p1 tile masks Preserve p1 tile masks
Branch Dense LoRA video Adaptive detail50 video Uniform 1×2 video
Control MP4 MP4 MP4
Preserve MP4 MP4 MP4

Prompt p2

The video is a simple yet elegant presentation of a green smoothie. The smoothie is served in a tall glass, which is placed on a white surface. The glass is adorned with a slice of orange and a sprig of mint, adding a touch of color and freshness to the scene. The background is a plain white wall, which contrasts with the green of the smoothie and the orange slice, making the drink the focal point of the image. The overall style of the video is minimalist and clean, with a focus on the smoothie and its ingredients. The video does not contain any text or additional elements, allowing the viewer to fully appreciate the smoothie and its presentation.

Control branch — step 500 Preserve branch — step 500
Control p2 tile masks Preserve p2 tile masks
Branch Dense LoRA video Adaptive detail50 video Uniform 1×2 video
Control MP4 MP4 MP4
Preserve MP4 MP4 MP4

Prompt p3

The video features a man in a suit, standing in a room with green walls. He is looking to his left, and his expression is serious. The room has a window in the background, and there is a blurred figure of a person in the foreground. The man is wearing a white shirt and a gray suit. The overall style of the video is realistic, with a focus on the man's expression and the room's interior.

Control branch — step 500 Preserve branch — step 500
Control p3 tile masks Preserve p3 tile masks
Branch Dense LoRA video Adaptive detail50 video Uniform 1×2 video
Control MP4 MP4 MP4
Preserve MP4 MP4 MP4

🟨 ENTROPY — best current candidate

Entropy was among the stronger sparse candidates in the initial layout screen: at step 0, entropy50's mean teacher gap was 0.429, compared with 0.416 for detail50 and 0.565 for uniform2 (lower is better). Detail50 was slightly better on that metric. This separate inference test compares matched 50% token-budget entropy placement with a root-shuffled entropy control at warm-up σ=0.9 and 30 steps. It uses four validation prompts only; it did not receive an untouched confirmation set and is not a promoted result.

ENTROPY is the best current candidate and the preferred algorithm for this project. The available selector run is a four-prompt validation screen, not an untouched confirmation. The initial teacher-gap screen slightly favored detail50 (0.416 vs. 0.429 for entropy50), so this preference should not be read as a statistically confirmed benchmark win.

The comparison preserves the original three sampled video frames under every method and appends that method’s matching t=0 tile grid as a fourth column. The gold outline marks the two ENTROPY rows. Each prompt sheet below shows all seven latent-time layouts alongside uniform and shuffled controls.

Three video frames per method followed by its matching tile grid

Prompt p0

The video features a young man posing confidently in front of a white background. He is shirtless, revealing a well-defined muscular physique, and has a tattoo on his left arm. He is wearing black shorts with white stripes on the sides. His smile is bright and he appears to be in a good mood. The video is likely a fitness or bodybuilding-related advertisement or promotional material. The style of the video is straightforward and focused on the man's physique, with no additional elements or distractions in the background.

ENTROPY p0 tile map and output frames

Prompt p1

The video features a person dressed in a vibrant costume, which includes a yellow top with orange sleeves and a pink wig with a bow on top. The individual is standing against a blue background, and their pose changes slightly between the frames. The style of the video is playful and colorful, with a focus on the costume and the person's expressive facial expressions. The overall mood of the video is cheerful and lighthearted.

ENTROPY p1 tile map and output frames

Prompt p2

The video features a man in a white lab coat and glasses, standing against a gray background. He appears to be a scientist or doctor, given his attire. The man is looking directly at the camera, suggesting that he is addressing the viewer. The lighting in the video is soft and even, highlighting the man's features without creating harsh shadows. The overall style of the video is professional and straightforward, with no additional elements or distractions. The focus is solely on the man and his message.

ENTROPY p2 tile map and output frames

Prompt p3

The video features a young boy with glasses, wearing a gray t-shirt. He is seated in a classroom setting, with a whiteboard and a bulletin board visible in the background. The boy appears to be engaged in a conversation or listening attentively. The style of the video is candid and informal, capturing a moment in the boy's daily life. The focus is on the boy, with the background elements providing context to the setting. The lighting is natural, suggesting an indoor environment with ample light. The video does not contain any text or additional graphics.

ENTROPY p3 tile map and output frames

Attribution

Base model: MiniMaxAI/MiniMax-H3, under the included H3 Community License and NOTICE. Training data and captions: OpenVid-1M / OpenVidHD, NJU-PCALab, CC BY 4.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for stale2000/LevelOfToken

Adapter
(126)
this model