Im running this on my 4070ti super (16 gb vram), and it takes 10 minutes for a 10-seconds 480p video. but the results are spectacular.
FWIW on a 5080 16GB it takes 3 minutes for 10 seconds 480p video (the mouse video workflow with length changed from 5 seconds to 10 seconds)
loading story #49161013
loading story #49156459
I am particularly curious how multimodal models will work with types of knowledge that are inherently non-text. For example, SOTA LLMs really suck at electronics, especially analog electronics.
Is MiniMax H3 capable of logical / technical reasoning, or is it purely art oriented?
loading story #49156717
loading story #49156580
loading story #49156674
loading story #49156598
1 minute for a second of footage, that is awesome! Thanks for sharing.
Huh, interesting. I tried to generate a 10 second 1080p clip on a bigger machine and the results were quite poor. Unusable for anything, in fact.
How much RAM does your machine have?