Hacker News new | past | comments | ask | show | jobs | submit

MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

https://blog.comfy.org/p/minimax-h3-day-0-support-in-comfyui
> We found that the model's modulation weights (~40% of the total parameters) could be pruned and replaced with a functionally equivalent lookup table, dramatically shrinking the memory footprint with no loss in output quality.

Is this a common approach to reducing weights with "no loss in output quality", assuming this is true? Seems almost too simple to work. If this is doable, would this be applicable to LLMs as well?

Neat with native frame-to-frame generation, but wonder how easy it is to "link" together clips at the intersection, typically the models kind of lose the "momentum" across these stiches, being able to merge things with frame-to-frame between clips might help with this it feels like.

loading story #49158433
loading story #49156519
loading story #49156407
loading story #49156860
Im running this on my 4070ti super (16 gb vram), and it takes 10 minutes for a 10-seconds 480p video. but the results are spectacular.
loading story #49156194
loading story #49156168
loading story #49156153
loading story #49156134
loading story #49156786
loading story #49156692
The mouse render is surprisingly good. Several of those clips stood out to be a pretty big leap in terms of current SOTA models.

The only one that looks "off" is the beverage ad video during the can opening clip, it still has that "AI smoothening" effect. Good thing this can be done pretty well using traditional rendering.

I feel like for a good while now we'll transition into a process that uses traditional "close-up" rendering/shots + AI generated wide-shots or quick cuts.

Exciting, but also troubling. This being open-weights is a massive win for the community though.

loading story #49156651
> The result gives a total memory footprint reduced by 66%, from 123.6 GB in full precision to 42.5 GB with the smallest models variants. Combining this with our dynamic VRAM offloading enables a next-generation 2K video model to run locally on a GPU like the RTX 3060.

Pretty cool.

But assuming you have a 16GB 3060, how long would it take to generate a 15 second clip?

On one hand: impressive. On the other hands aesthetically it all looks painfully bland and generic.
loading story #49156469
loading story #49156214
loading story #49160438
Reference-to-video mode seems like all that was missing to enable completely independent cinematography as right now one couldn't stitch different scenes together properly without altering substantial portions of the scene.
I saw the samples people have posted. Immediately deleted LTX2 and WAN folders. Those are completely worthless now.

There is some debate on the license for those in the US, UK, EU, plus… no comment other than whew those samples though!

loading story #49156089
I've said it before and I'll say it again, human directors are still valuable, as they use AI video editing tools to generate the shots they want and put them together in a cohesive way. Previously they might've used film and actors but if they can just prompt the AI (or create workflows as seen with ComfyUI) then they arrange them together just like how an EDM producer doesn't actually play the instruments but instead the creativity is in the arrangement.

I suspect it'll be quite a while until AI gets a good enough aesthetic sense to do this, as even with static HTML websites humans can easily see that it's AI slop.

loading story #49157433
Learned something, upvoted
Any tutorial for me to learn how to use
loading story #49160590
Hollywood and the film industry on red alert. Too bad.

This is AGI.

The example video just looks like the highly produced art (TV, commercials, games) other people have created. I find it impossible to believe this wasn't trained on other people's work, and there is no protection for it. Terribly sad. A lack of original thinking is coming.
loading story #49156457
loading story #49157743
loading story #49155932
Unlimited slop "content" to fill decomposed brain shaped vessels, the future is bright!
loading story #49156237