MiniMax-H3 Level-of-Token research release
Experimental LoT adapter and rank-16 DiT LoRA, with benchmarks and visual examples.
Research status: the latest campaign found no branch that met the predeclared promotion threshold. Metrics show tradeoffs, not a confirmed quality win. See benchmark results and campaign summary.
What was tried
We measured token-layout speed, screened sparse layouts, trained on still images and OpenVidHD clips, compared a control against dense-teacher preservation, and tested inference-time detail/entropy selectors. The best measured DiT-forward speedup was 2.64× on a 37-frame 768×1344 case; VAE decode was unchanged. The r4 500-update branches were effectively tied on the 24-source confirmation set and neither passed promotion criteria. Selector previews were small validation screens, not confirmation benchmarks. Read the experiment history and values.
Weights
adapter.safetensors: LoT adapter and fitted extent bank.lora.safetensors: rank-16 DiT LoRA.
Four-column video comparisons
Both grids show matched 30-step outputs from the untouched confirmation prompts. The fourth column shows the matching token layout for each row (latent t=0); source-reference rows are marked as having no model tile mask. Full seven-frame layouts appear in the prompt sheets below. The downloadable MP4 links are generated outputs.
Control branch, step 500
Preserve branch, step 500
Prompts, videos, and tile-layout masks
The mask sheets show all seven latent-time layouts for each prompt. Each sheet compares dense 1×1, uniform 1×2, and the adaptive detail50 pattern recorded at the warm-up selection step. These are token-layout maps, not semantic segmentation masks. Each prompt has separate masks for the control and preserve checkpoints.
Prompt p0
In the video, a man wearing glasses and a blue shirt is seen interacting with a small dog. The dog, with its white and brown fur, is sitting on a pink surface. The man appears to be speaking to the dog, possibly giving it a command or a treat. The scene is set against a white wall, which provides a neutral backdrop for the interaction between the man and the dog. The overall style of the video is casual and intimate, capturing a moment of connection between the man and his pet.
| Branch | Dense LoRA video | Adaptive detail50 video | Uniform 1×2 video |
|---|---|---|---|
| Control | MP4 | MP4 | MP4 |
| Preserve | MP4 | MP4 | MP4 |
Prompt p1
In the video, a man is seen interacting with a yellow and black device attached to a green door. The device appears to be a camera or a sensor, and the man is holding it up to the door, possibly taking a picture or recording a video. The man is dressed in a black jacket and seems to be focused on his task. The door is green and has a window, through which the man's reflection can be seen. The scene suggests that the man might be a security officer or a technician, inspecting or setting up the device for security purposes. The overall style of the video is realistic and straightforward, with no additional elements or distractions.
| Branch | Dense LoRA video | Adaptive detail50 video | Uniform 1×2 video |
|---|---|---|---|
| Control | MP4 | MP4 | MP4 |
| Preserve | MP4 | MP4 | MP4 |
Prompt p2
The video is a simple yet elegant presentation of a green smoothie. The smoothie is served in a tall glass, which is placed on a white surface. The glass is adorned with a slice of orange and a sprig of mint, adding a touch of color and freshness to the scene. The background is a plain white wall, which contrasts with the green of the smoothie and the orange slice, making the drink the focal point of the image. The overall style of the video is minimalist and clean, with a focus on the smoothie and its ingredients. The video does not contain any text or additional elements, allowing the viewer to fully appreciate the smoothie and its presentation.
| Branch | Dense LoRA video | Adaptive detail50 video | Uniform 1×2 video |
|---|---|---|---|
| Control | MP4 | MP4 | MP4 |
| Preserve | MP4 | MP4 | MP4 |
Prompt p3
The video features a man in a suit, standing in a room with green walls. He is looking to his left, and his expression is serious. The room has a window in the background, and there is a blurred figure of a person in the foreground. The man is wearing a white shirt and a gray suit. The overall style of the video is realistic, with a focus on the man's expression and the room's interior.
| Branch | Dense LoRA video | Adaptive detail50 video | Uniform 1×2 video |
|---|---|---|---|
| Control | MP4 | MP4 | MP4 |
| Preserve | MP4 | MP4 | MP4 |
🟨 ENTROPY — best current candidate
Entropy was among the stronger sparse candidates in the initial layout screen: at step 0, entropy50's mean teacher gap was 0.429, compared with 0.416 for detail50 and 0.565 for uniform2 (lower is better). Detail50 was slightly better on that metric. This separate inference test compares matched 50% token-budget entropy placement with a root-shuffled entropy control at warm-up σ=0.9 and 30 steps. It uses four validation prompts only; it did not receive an untouched confirmation set and is not a promoted result.
ENTROPY is the best current candidate and the preferred algorithm for this project. The available selector run is a four-prompt validation screen, not an untouched confirmation. The initial teacher-gap screen slightly favored detail50 (0.416 vs. 0.429 for entropy50), so this preference should not be read as a statistically confirmed benchmark win.
The comparison preserves the original three sampled video frames under every method and appends that method’s matching t=0 tile grid as a fourth column. The gold outline marks the two ENTROPY rows. Each prompt sheet below shows all seven latent-time layouts alongside uniform and shuffled controls.
Prompt p0
The video features a young man posing confidently in front of a white background. He is shirtless, revealing a well-defined muscular physique, and has a tattoo on his left arm. He is wearing black shorts with white stripes on the sides. His smile is bright and he appears to be in a good mood. The video is likely a fitness or bodybuilding-related advertisement or promotional material. The style of the video is straightforward and focused on the man's physique, with no additional elements or distractions in the background.
Prompt p1
The video features a person dressed in a vibrant costume, which includes a yellow top with orange sleeves and a pink wig with a bow on top. The individual is standing against a blue background, and their pose changes slightly between the frames. The style of the video is playful and colorful, with a focus on the costume and the person's expressive facial expressions. The overall mood of the video is cheerful and lighthearted.
Prompt p2
The video features a man in a white lab coat and glasses, standing against a gray background. He appears to be a scientist or doctor, given his attire. The man is looking directly at the camera, suggesting that he is addressing the viewer. The lighting in the video is soft and even, highlighting the man's features without creating harsh shadows. The overall style of the video is professional and straightforward, with no additional elements or distractions. The focus is solely on the man and his message.
Prompt p3
The video features a young boy with glasses, wearing a gray t-shirt. He is seated in a classroom setting, with a whiteboard and a bulletin board visible in the background. The boy appears to be engaged in a conversation or listening attentively. The style of the video is candid and informal, capturing a moment in the boy's daily life. The focus is on the boy, with the background elements providing context to the setting. The lighting is natural, suggesting an indoor environment with ample light. The video does not contain any text or additional graphics.
Attribution
Base model: MiniMaxAI/MiniMax-H3, under the included H3 Community License and NOTICE. Training data and captions: OpenVid-1M / OpenVidHD, NJU-PCALab, CC BY 4.0.
Model tree for stale2000/LevelOfToken
Base model
MiniMaxAI/MiniMax-H3













