I’ve been loving all the new nodes and workflows coming out for MinMax, and maybe there is already a nice solution for this - but I couldn’t find one that did exactly what I needed.
I started using MinMax for my last TBG ETUR video and quickly ran into limitations: I wanted an easy way to create lip-sync videos longer than 20 seconds.
I didn’t want to manually chain ComfyUI nodes, start a new run every X seconds, or constantly resize things just to make HD video fit into my available VRAM.
The addon automatically chains MinMax H3 lip-sync generations together, allowing you to create much longer lip-sync videos without manually setting up each 20-second segment.
The workflow has a simple switcher that lets you switch from the 32B CLIP to the 4B CLIP, saving around 10 GB of VRAM. You can also switch from Sage to Comfy Kitchen, Spectrum to Easy Cache, or FL2VA to REF2VA both setup for lip-syncing. Some of it could be useful for other tasks as well.
You will find the workflow in the repro and tested recommendations, optimized settings, presets, and more workflows, along with the results of my testing and performance here
Small critique: she needs to remember to “breathe”. The longer the video goes on, the more the sentences were just crashing together, without pause,kindoflikereadingawholrthing justlikethis withouttheproper timing and breaks.
While the presenter is offscreen, slow it down a notch, break the sentences with an added half second pause, you can easily cheat a more natural video without breaking the existing lypcsyn, and still try and create pacing waves.
It would be a post-process edit you’d have to make using audio software, adjust video speeds with something like Premiere, splice things back together.
Just because it was output at a specific rate doesn’t mean you’re limited to that one clip!
Yep — it was outside my timeline for this post, but you’re right. For production videos where clients and money are involved, this definitely needs to be done. However, this is just an intro and tutorial for a free node pack, so I don’t have the same criteria to apply here.
Thank you for sharing the work and video. Looks very powerful and useful. What did it take to make this video? I think I spotted a few of those reworked segments mentioned in the clip. Do you manually run each segment, repeating when needed, or something else?
What I usually do is set up the workflow, run it 3–4 times, and then combine the best shots afterward.
Some of the glitches you see come from runs I did with the Turbo LoRA, including things like degraded skin. So the examples are a mix of different testing settings and the final setup.
The final settings and node should work out of the box; the main factor is the prompt. For lip-sync, I recommend not using Turbo and sticking to 20+ steps, along with a good, detailed prompt.
Very interesting.
Does this only work with a audio file input.
Or could you also use a prompt for a let’s say 2 minute sequence and split that up?
That’s what’s puzzling me rn.
Generation a very long shot list. Then divide that into chunks and feed it to separate samplers. But difficulty is the possible change of subjects etc settings.
So now I separate a master prompt that gets written by one h3 prompt writer node, separate out the shot list and divide that into chunks.
Then put it back together with the stuff that comes before shots so each sampler gets the prompt plus the specific shot list for that chunk. But it’s still not working well. Because the prompt is for the whole story and then the shot list lacks the context of what came before.
So it can get messed up.
Also the whole prompt writing thing in itself sometimes fails to get what are the subjects and what to do with what.
Feeding the separate chunk timings to separate Minimax h3 prompt weiter nodes frequently gets stuff wrong.
A workflow to take one giant shortlist and reliably correctly feed that to the samplers or render it is what I would like to achieve.
The model accepts prompts like “Subject 1 says, ‘Too hard for me.’” But someone would have to split the prompt into individual clip proportion and cuting at the exact final second for each one, which would be complicated. It’s much easier to pass the text directly to a TTS system first. Feeding them into separate samplers isn’t possible either, because each clip depends on the previous one.
Man, I’m gonna kill myself for idea to make a fan ai video clip on a song. 2 days passed and I made just a minute or so from that video, spending time for 2 generations 8 sec each, first with turbo Lora to see if prompt was good, second without Lora with 25+ steps. I babysitted each 8seconds clip, finding ideas, using references, and now you tell me a can do it in one run? Will try it asap, I love you
This is absolute GOLD! I am in the process of playing with MiniMax H3 and I am definitely saving this to reference as I get into the process. Thank you so much!
Impressive. The better these workflows get the more nitpicking wants to happen though. Long sleeve vs tsshirt? Sudden gradient background? Mic switching sides? It becomes very distracting. Uncanny almost.
😄 There’s a lot to fight with - prompts and conditioning. My approach is to let it generate around X videos at different clip lengths, then pick the best results from all of them - only really necessary if you need it to be as perfect as possible. This video was made from the “garbage” I had left over from building and testing the node.
Yes you just need to find the right prompt. You can also try using ref 0 in the style section of the prompt to keep the camera position consistent for each frame. You might use an image without the character as a reference for this and ref 1 only for the character. You’ll have to test a few variations to see what works best. Its all about the right promt.
You need to define the clip length yourself based on your VRAM limit. The node handles the rest: it cuts the audio in clips, creates the overlaps, and adds 24 frames of silence at the end to improve the final image and sound. It then uses the latent from each clip to build the appropriate motion for the next one and so on.
Let's say if I have a 10min audio file. If I set the clip length as 10sec. Then it will generate a 10min video synced with the audio with 10sec clips stitched together automatically?
I would first build the full video with the audio in an editor, and then only generate/sample the 2–3 minute sections I actually need for the final production without cuts or switching to non-character scenes.
That makes it much easier to prompt and much faster to rerender or repeat individual sections than trying to generate 10 minutes in one go.
Anyway each clip segment gets its own output in the output folder, so you can stop and resume from the clips you already have, or simply rerun one specific clip later. (start end inputs)
The node will recognize the existing clips and automatically rebuild the new full-length clip with the repeated/replaced segment included of the same id.
Out of curiosity, and for someone with low experience with these types of workflows, can you share what hardware you used and how long it took to get to this end result? I know you said in another comment that this was a bunch of garbage clips stitched, but trying to gauge how long something like this actually takes.
And for those of us that don't have the hardware - any recommendations for trying these things out in ComfyUI Cloud (or something else)?
20 steps at 20 seconds each, with 1 MP resolution, take about 50 seconds of sampling per step, so the full run takes around 16 minutes 40 seconds.
To speed things up, I start with 0.3–0.4 MP, which brings the sampling time down to around 9 seconds per step. With an efficient cache, this can be reduced by roughly half, to around 4–5 seconds per step.
This lets me quickly check whether the prompt is good. If it is, I then run the longer, high-resolution version.
I needed one 60-second video, one 100-second video, and two 32-second videos. Let’s say I generated each one 3×, which comes to:
* 60 sec × 3 = 180 sec
* 100 sec × 3 = 300 sec
* 32 sec × 3 × 2 = 192 sec
* Corrections: 5 × 10 sec = 50 sec
Total: 722 seconds, or about 12 minutes.
Sampling over night:
If 1 second of final video takes 50 seconds to sample/generate, then:
12 min video → 722 sec → × 50 sec/second = 36,100 sec ≈ 10 hours
I gave myself a day for it so I did everything in one day: the video, text, TTS, node programming, MinMax research, and workflow comparison.
It’s VibeVoice itself that sometimes adds the audio. I was too lazy to run it again, not H3—the audio is entirely from the input. I just took the image part from the generated minmax H3 video.
Ah, okay. I got a weird thing, where I listen behind people speaking for the weird background audio -- my favourite for this is the TV show 'Corner Gas', which has really complex background audio.
Yeah, it's fine, some of them were just quite noticable.
I wonder if there's a package out there for filtering that kind of thing out: if you have the audio and the transcript, which we do, it seems possible for a machine to clean up. Bound to be. Seems like something that would exist.
Much easier - you just change the seed in VibeVoice, and after 1–2 tries you usually get a version without the unwanted background sound. And yes, there are models that can filter out the voice and clean up the audio for comfyui
Ehhh... nah. Reroll is the lazy man's way, and it probably costs more.
The point of a filter is that you can apply it to any output: even one that is already clean. So, you can automate the whole flow, and not really have to worry about non-deterministic problems.
The input image was from a very old Flux.1 pic. The bad grid came from H3 + Turbo with too few steps and, I think, overly strong caching. I probably kept one or two cuts from that runs because the good final run had some strange head or hand movements.
I think so. I’ve never done it myself, but the model should understand it. I’ve seen workflows that send batches of images as references, so it’s up to you to test. The references you use are also up to you. My node only changes ref 0 for each clip; all the other references are passed through to every clip.
Can somebody just tell me the basics of how this audio sounds so damned clean and why others sound so "corrugated" and dithered? Does it require an audio source, or can it be natural sounding from t2v if you make precise character/actor references with the right workflow/inference settings?
The audio source is a high-quality live recording with similar tutorial speech. I had to test several of them before finding the one that worked for me.
Thank you for sharing. I cannot install your nodes pack: I copied the files into custom_nodes in ComfyUI-H3-Motion-Context-Auto-Chain folder, but nodes are still unavailable. What do I do wrong? It is possible to install this pack via ComfyUI Manager? Right now the Manager says Node 'comfyui-h3-motion-context-auto-chain-addon@nightly' not found in [default, cache]
Looked through the ComfyUI startup log and see the following error message:
It all depends on your hardware and settings. On a 5090, I can do a 30-second generation at 1MP with 20 steps. Fewer steps, a smaller image size, and more VRAM will all affect how many secs you can generate.
I happen to also have a 5090 ( literally the best I have and could afford ). I did try once Minimax but 15 seconds was a nightmare takes forever? May I ask how long does on average take 30 sec generations for you so I can atleast have a point of reference? Thanks, cool video too!
Using my WF and the 5090 power limit at 85%, fans at 100% (mem 15,891 MHz, GPU 2,377 MHz, undervolted to 869 mV = max 2,500 MHz, max temp limit 82°C).
Only for ComfyUI, not as a display, with 4B CLIP and 20-30 steps, Easy Cache on + Kitchen on Results 5 sec | 30 steps | 0.4 MP: 26.5 GB VRAM used | ~70 sec total 5 sec | 30 steps | 1 MP: 28.9 GB VRAM used | 8 it/s | ~180 sec total (Before someone clever asks: it’s not 8×30 because EasyCache significantly reduces the processing time.) 15 sec | 20 steps | 1 MP: 24.6 GB VRAM used | 51 it/s | 714 sec total | CLIP on RAM 15 sec | 20 steps | 1 MP: 25.4 GB VRAM used | 51 it/s | 714 sec total | 4B CLIP on VRAM 20 sec | 20 steps | 1 MP: 21 GB VRAM used | 81 it/s | 1134sec total 30 sec | 20 steps | 1 MP: OOM | --disable-dynamic-vram --disable-pinned-memory | without these options it runs even slower on RAM, which I think is your case. 28 sec | — steps | 1 MP: 28.3 GB VRAM used | 145.85 it/s | 2381 sec total | very slow 27 sec | 20 steps | 1 MP: 27.6 GB VRAM used | 140 it/s | 1960 sec total 30 sec | 20 steps | 1 MP: 30.3 GB VRAM used | 166 it/s | 2324 sec total | CLIP on RAM to prevent OOM
More seconds = more frames → a larger temporal latent → more temporal tokens for the Transformer to process at every denoising step. Therefore, the GPU has to calculate more latent data per step, so longer videos require disproportionately more GPU compute.
So if your goal is 30 seconds of continuous video, the interesting comparison is if 30s single latent vs 6×5s independent clips is better ? — Try 30s single latent vs 6×5s with proper H3 latent/context continuation. The latent can preserve continuity while avoiding the huge nonlinear cost you’re seeing.
To answer your question about how long it takes to generate 30 seconds: it depends on the approach. 6 × 5 sec = ~15 minutes total with proper H3 latent/context continuation 1 × 30 sec = ~60 minutes total
So generating 30 seconds continuously takes about 4× longer than generating six 5-second clips.
Thanks so much!! This is so detailed. Also I had no clue there was a proper way for H3 latent/context continuation, I’ll definitely look into that too. Thanks again 🙏!!
try latest version from github or comfyui manager and use the new workflows - than enable endless_continuation.
This is not using anymore last frame as the first frame. The degradation may be coming from the last frame because, depending on your settings, the quality can decrease toward the end.
Try using a resolution of around 2–2.5 MP. The lip-sync and overall details especially things like hand movement are much better at higher resolution than at low resolution.
Legend, thanks.
I’ve got a list of diff models to test that has VibeVoice in from a while back..
I’d just been using video models to get short clips but voice continuity has been annoying me haha
Not exactly true — it was built from different runs while I was building and testing the node. The examples are a mix of 8–40 steps, both with and without Turbo LoRA, and both with and without caching.
So the quality varies depending on which final cuts I selected. The workflow does produce a full-length video, but I didn’t use just one single run. Since this is AI, if I had the wrong clip prompt and, for example, the character was just listening instead of speaking, I would rerender that individual clip and continue from there. That’s pretty normal with AI workflows — it’s rarely just one run from start to finish.
If you set it to a fixed 20+ steps without Turbo, the model will hold up well.
I get same error :
Required input missing
H3 Auto Chain Motion Context is missing a required input: context_frames
[ERROR] Failed to validate prompt for output 426:418:
[ERROR] * MiniMaxH3AutoChainMotionContext 426:423:
[ERROR] - Required input is missing: context_frames
[ERROR] Output will be ignored
[ERROR] Failed to validate prompt for output 293:
[ERROR] Output will be ignore
[ERROR] Failed to validate prompt for output 426:417:
[ERROR] Output will be ignored
[WARNING] invalid prompt: {'type': 'prompt_outputs_failed_validation', 'message': 'Prompt outputs failed validation', 'details': '', 'extra_info': {}}
The addon inherits its input definition from ComfyUI-H3-Motion-Context. If you have an older or modified version installed, you may encounter this error.
Alternatively, update my node to v0.1.2. I’ve made the node independent of ComfyUI-H3-Motion-Context, so future updates or changes to ComfyUI-H3-Motion-Context won’t affect my node. THIRD_PARTY_NOTICES.md is included as well.
v0.1.1 remains the release version that depends on ComfyUI-H3-Motion-Context.
31
u/mfdi_ 8d ago
Apart from ai looking character. Wow. Just wow. Though camera moving kinda sucks.