Hacker News new | past | comments | ask | show | jobs | submit
I wonder how much of that is due to the use of a suboptimal VAE? The same optimisations could be applied to it, and my intuition tells me that your total compute spend would be even more optimal if you retrained the VAE yourself with a better approach (esp. ensuring translation and rotation invariance, + ability to rescale/blur the latents).

You could also have the same advantages of a draft in low resolution latent space, with an easier to learn data distribution that's more robust to perturbation, but instead of 512x512, you could get a 4096x4096 output for the same compute (assuming a 8x VAE).

There's also the advantage of being able to use a high number of diffusion/flow matching steps for the main LDM while the VAE can be single step and much smaller, since it does not have to handle language or significant scene understanding, just perceptual compression. This sounds especially important for a video model where I would be extremely hesitant to train a generative model without relying on interframe compression.

loading story #50025984
loading story #50025847