This is what happens when you optimize the model and inference stack together. We post-trained H3 Max for stronger quality and prompt adherence, then pushed inference until a 5-second video could generate in under 3 seconds, faster than real time.
It feels like a big line in AI video was crossed, in my experiments, H3 Max can now create reasonably high quality AI video in less time than it takes you to watch it. This is realtime from the moment I pushed the "generate" button (and also includes the time for it to improve my prompt about an otter astronaut automatically)